All language subtitles for AI And Machine Learning Full Course [FREE] Learn AI And Machine Learning In 24 Hours Simplilearn.English

af Afrikaans
ak Akan
sq Albanian
am Amharic
ar Arabic
hy Armenian
az Azerbaijani
eu Basque
be Belarusian
bem Bemba
bn Bengali
bh Bihari
bs Bosnian
br Breton
bg Bulgarian
km Cambodian
ca Catalan
ceb Cebuano
chr Cherokee
ny Chichewa
zh-CN Chinese (Simplified)
zh-TW Chinese (Traditional)
co Corsican
hr Croatian
cs Czech
da Danish
nl Dutch
en English
eo Esperanto
et Estonian
ee Ewe
fo Faroese
tl Filipino
fi Finnish
fr French
fy Frisian
gaa Ga
gl Galician
ka Georgian
de German
el Greek
gn Guarani
gu Gujarati
ht Haitian Creole
ha Hausa
haw Hawaiian
iw Hebrew
hi Hindi
hmn Hmong
hu Hungarian
is Icelandic
ig Igbo
id Indonesian
ia Interlingua
ga Irish
it Italian
ja Japanese
jw Javanese
kn Kannada
kk Kazakh
rw Kinyarwanda
rn Kirundi
kg Kongo
ko Korean
kri Krio (Sierra Leone)
ku Kurdish
ckb Kurdish (Soranรฎ)
ky Kyrgyz
lo Laothian
la Latin
lv Latvian
ln Lingala
lt Lithuanian
loz Lozi
lg Luganda
ach Luo
lb Luxembourgish
mk Macedonian
mg Malagasy
ms Malay
ml Malayalam
mt Maltese
mi Maori
mr Marathi
mfe Mauritian Creole
mo Moldavian
mn Mongolian
my Myanmar (Burmese)
sr-ME Montenegrin
ne Nepali
pcm Nigerian Pidgin
nso Northern Sotho
no Norwegian
nn Norwegian (Nynorsk)
oc Occitan
or Oriya
om Oromo
ps Pashto
fa Persian
pl Polish
pt-BR Portuguese (Brazil)
pt Portuguese (Portugal)
pa Punjabi
qu Quechua
rm Romansh
nyn Runyakitara
ru Russian
sm Samoan
gd Scots Gaelic
sr Serbian
sh Serbo-Croatian
st Sesotho
tn Setswana
crs Seychellois Creole
sn Shona
sd Sindhi
si Sinhalese
sk Slovak
sl Slovenian
so Somali
es Spanish
es-419 Spanish (Latin American)
su Sundanese
sw Swahili
sv Swedish
tg Tajik
ta Tamil
tt Tatar
te Telugu
th Thai
ti Tigrinya
to Tonga
lua Tshiluba
tum Tumbuka
tr Turkish
tk Turkmen
tw Twi
ug Uighur
uk Ukrainian
ur Urdu
uz Uzbek
vi Vietnamese
cy Welsh
wo Wolof
xh Xhosa
yi Yiddish
yo Yoruba
zu Zulu

Original subtitles

Welcome to simply learns YouTube

channel. Artificial intelligence and

machine learning are transforming the

way business operate, make decisions and

innovate. From personalized

recommendation on streaming platforms

and intelligent chat bots to

self-driving vehicles and advanced

healthcare systems, AI and machine

learning are powering some of the most

impactful technologies of our time. And

AI and machine learning engineering

combines programming, mathematics,

statistics, data science, machine

learning expertise to create intelligent

applications that deliver business

value. In this complete AI and machine

learning engineering course, you will

learn everything from the fundamentals

of AI and machine learning to advanced

concept used in the modern intelligent

systems. We will start with the

programming and mathematical foundation

then gradually move into data analysis,

machine learning algorithms, deep

learning and generative AI. Throughout

the course, you will gain hands-on

experience with industry standard tools

and frameworks such as Python, NumPy,

Pandas, Kikit Learn, TensorFlow,

PyTorch, and others. You'll also learn

how to collect and prepare data, build

predictive models, train neural

networks, evaluate model performance,

and deploy machine learning solution in

real world environments. By the end of

this course, you'll have a strong

understanding of complete AI and machine

learning life cycle and practical skills

required to pursue a career as a AI and

machine learning engineer. Having said

that, let's take a look at today's

agenda. We'll start off with module one,

which is introduction to artificial

intelligence and machine learning.

Module two is Python programming for AI

and ML. Module three is mathematics,

statistics, and probability for machine

learning. Module four is data

collection, cleaning, and

pre-processing. Module five is

exploratory data analysis and data

visualization. Module six is machine

learning fundamentals. Module seven is

supervised learning algorithms. Module

eight is unsupervised learning

algorithms. Module 9 is model evaluation

and feature engineering. Module 10 is

deep learning and neural networks.

Module 11 is natural language

processing. Module 12 is computer vision

fundamentals. Module 13 is generative

AI, LLMS and AI agents. Module 14 is

envelopes, model deployment and AI

engineering workflows. Module 15 is real

world AI projects. Module 16 is

interview question and answers. Hope I

made myself clear with that agenda.

That's it. If these are the type of

videos you would like to watch, then hit

that subscribe button with the bell icon

to get notified whenever we host. Also,

just so that you know, if you want to

upskill yourself, master generative AI

and land your dream job or even grow in

your career, then you must explore

Simply Learn's cohort of various

generative AI training and professional

certification programs. Simply learn

offers a variety of masters

certification and post-graduate programs

in collaboration with some of the

world's leading universities. Through

our courses, you will gain knowledge

along with work ready expertise in

skills like Python, Aentic AI, AI

automation systems, LLMs, and over a

dozen others. And that's not all. You

will also get the opportunity to work on

multiple projects led by industry

experts working on top tier

service-based and product companies.

After completing these courses,

thousands of learners have transition

into an AI and machine learning role as

a fresher or moved on to your higher

paying job and profile. If you're

passionate about making your career in

this field, then make sure to check out

the link in the pinned comments and in

the description box to find an AI and

machine learning program that fits your

experience and areas of interest. So

let's get started with our AI and

machine learning engineer full course

with a small quiz. What is NLP? Is it

network layer processing, natural

language processing, neural learning

platform, or is it native language

programming? Please let us know your

answers in the comment section below.

Now over to our training experts.

>> Once upon a time in the quiet town of

Newite, there lived a curious teenager

named Arya. She wasn't like most kids in

her school. While others were busy with

sports or music, Arya was fascinated by

machines, especially the idea of making

machines think like humans. Her

curiosity began one evening when she

asked her grandfather, who used to be a

computer engineer, "Can machines ever

think?" Her grandfather smiled and said,

"That's what artificial intelligence is

all about." Arya's eyes lit up.

Artificial intelligence? What's that? So

he began to tell her a story, not a

fairy tale, but a real story about the

science and ideas behind machines that

learn, decide, and sometimes even

surprise their creators. Artificial

intelligence, or AI, is the science of

making machines that can do things that

normally require human intelligence.

This includes tasks like recognizing

faces, understanding speech, making

decisions, and even playing games. But

AI isn't magic. It's built through

programming, mathematics, and data. Arya

imagined a robot that could talk like a

human and help with homework. Her

grandfather nodded. That's one kind of

AI, but there are many types. He

explained that AI isn't just about

robots. In fact, most AI systems are

just computer programs running inside

machines we already use, like phones,

laptops, or even refrigerators. Her

grandfather told her that AI comes in

two main types, narrow AI and general

AI. Narrow AI is the kind we see today.

It's designed to do one specific task.

For example, the AI in a smartphone that

unlocks the screen by recognizing your

face is only good at that one job. It

can't cook or write a story. General AI,

on the other hand, would be as smart as

a human, able to learn anything and do

many tasks. But this type of AI doesn't

exist yet. It's more of a dream for now.

Arya asked, "How do these machines

learn? That's where machine learning

comes in," her grandfather replied.

Machine learning is a type of AI that

learns from data instead of being told

what to do step by step. "Imagine

teaching a dog to sit. You show it how,

give it treats, and repeat. Over time,

the dog learns. Machine learning works

the same way. You feed it data and it

finds patterns. For example, if you want

a computer to recognize pictures of

cats, you show it thousands of cat

pictures. It starts to see what cats

usually look like. Furry whiskers,

pointy ears. Over time, it learns to

tell a cat apart from a dog or a chair.

The program that does this learning is

called a model. A model is like a brain

built by the computer using the data it

was given. The more data it gets, the

better it learns. But how does the

computer know what a cat is? Arya asked.

Her grandfather said, "That's thanks to

something called a neural network. It's

a method used in machine learning that's

inspired by how our brains work. A

neural network is made up of layers of

tiny parts called neurons. These are not

real brain cells, but math functions.

Each neuron takes in numbers, does some

math, and passes the result to the next

layer of neurons. Imagine passing a note

through a group of friends, and each one

adds or changes a word before giving it

to the next. By the end, the note may

have transformed in a useful way. That's

what a neural network does to data. It

turns it into something meaningful, like

recognizing a cat in a picture. The more

layers a network has, the more complex

patterns it can understand. When a

network has many layers, it's called

deep learning. To get a neural network

to work, it needs to be trained.

Training is the process where the model

is shown lots of examples so it can

learn. Training involves giving the

model data and letting it guess

something like whether a picture has a

cat. At first, it guessed badly, but

then it compares its guess to the

correct answer. If it's wrong, it

adjusts itself using a method called

back propagation. Back propagation is

like checking your math homework. If the

answer is wrong, you go back, find where

you messed up, and fix it. In AI, this

helps the model improve step by step.

This cycle of guessing, checking, and

adjusting is repeated many times. The

model slowly gets better at the task.

Can AI make mistakes? Arya asked. Oh

yes, her grandfather said AI is smart in

some ways but not perfect. AI only

learns from the data we give it. If the

data is bad, the AI will be bad. This is

called bias. For example, if a face

recognition system is trained mostly on

photos of light-kinned people, it might

not work well on darkerkinned people.

Also, AI doesn't really understand the

world. It only sees patterns in numbers.

It doesn't know what a cat feels like or

why we love them. That's why AI can

sometimes be fooled by simple tricks

like weird images that a human would

never mistake for a cat. AI is

everywhere, her grandfather explained.

It helps recommend videos on YouTube,

powers voice assistants like Siri or

Alexa, drives some cars, and even helps

doctors find diseases and scans. But not

all AI is harmless. It can be used for

spying, spreading fake news, or making

decisions that affect people's lives,

like who gets a loan or a job? That's

why it's important for people to

understand how AI works so they can ask

good questions and build it responsibly.

Arya asked, "Will AI take over the

world?" Her grandfather laughed. Not

like in the movies, but it will change

the world. The future of AI depends on

how people choose to use it. It can help

solve big problems like climate change

or disease. But it also needs rules and

careful thinking. Just like fire or

electricity, AI is a tool, a powerful

one. If used wisely, it can do great

good. Arya sat back, her mind buzzing.

She had started the day wondering if

machines could think. Now she knows that

while they don't think like humans, they

can do amazing things through learning

data, and clever programming. She smiled

and said, "Maybe I'll build an AI

someday." Her grandfather smiled, too.

Just remember, it's not about making a

machine smart. It's about making it

useful and fair for everyone. And from

that day on, Arya started her journey

not just to understand AI, but to shape

it with care, creativity, and curiosity.

>> We know humans learn from their past

experiences, and machines follow

instructions given by humans.

But what if humans can train the

machines to learn from their past data

and do what humans can do and much

faster? Well, that's called machine

learning. But it's a lot more than just

learning. It's also about understanding

and reasoning. So today we will learn

about the basics of machine learning. So

that's Paul. He loves listening to new

songs.

He either likes them or dislikes them.

Paul decides this on the basis of the

song's tempo, genre, intensity, and the

gender of voice. For simplicity, let's

just use tempo and intensity for now.

So, here tempo is on the x-axis, ranging

from relaxed to fast, whereas intensity

is on the y-axis, ranging from light to

soaring. We see that Paul likes the song

with fast tempo and soaring intensity

while he dislikes the song with relaxed

tempo and light intensity. So now we

know Paul's choices. Let's say Paul

listens to a new song. Let's name it as

song A. Song A has fast tempo and a

soaring intensity. So it lies somewhere

here. Looking at the data, can you guess

whether Paul will like the song or not?

Correct. So Paul likes this song. By

looking at Paul's past choices, we were

able to classify the unknown song very

easily, right? Let's say now Paul

listens to a new song. Let's label it as

song B. So song B lies somewhere here

with medium tempo and medium intensity.

Neither relaxed nor fast, neither light

nor soaring. Now, can you guess whether

Paul likes it or not? Not able to guess

whether Paul will like it or dislike it.

Are the choices unclear? Correct. We

could easily classify song A. But when

the choice became complicated as in the

case of song B. Yes. And that's where

machine learning comes in. Let's see

how. In the same example for song B, if

we draw a circle around the song B, we

see that there are four votes for like

whereas one vote for dislike. If we go

for the majority votes, we can say that

Paul will definitely like the song.

That's all. This was a basic machine

learning algorithm also. It's called K

nearest neighbors. So this is just a

small example in one of the many machine

learning algorithms quite easy right

believe me it is but what happens when

the choices become complicated as in the

case of song B that's when machine

learning comes in it learns the data

builds the prediction model and when the

new data point comes in it can easily

predict for it more the data better the

model higher will be the accuracy there

are many ways in which the machine

learns it could be either supervised

learning unsupervised learning or

reinforcement learning. Let's first

quickly understand supervised learning.

Suppose your friend gives you 1 million

coins of three different currencies. Say

1 rupee, 1 and 1 dirham. Each coin has

different weights. For example, a coin

of 1 rupee weighs 3 g. 1 euro weighs 7 g

and 1 dirham weighs 4 g. Your model will

predict the currency of the coin. Here

your weight becomes the feature of coins

while currency becomes their label. When

you feed this data to the machine

learning model, it learns which feature

is associated with which label. For

example, it will learn that if a coin is

of 3 g, it will be a 1 rupee coin. Let's

give a new coin to the machine. On the

basis of the weight of the new coin,

your model will predict the currency.

Hence, supervised learning uses labeled

data to train the model. Here, the

machine knew the features of the object

and also the labels associated with

those features. On this note, let's move

to unsupervised learning and see the

difference. Suppose you have cricket

data set of various players with their

respective scores and the wickets taken.

When we feed this data set to the

machine, the machine identifies the

pattern of player performance. So, it

plots this data with the respective

wickets on the x-axis while runs on the

y-axis. While looking at the data,

you'll clearly see that there are two

clusters. The one cluster are the

players who scored high runs and took

less wickets while the other cluster is

of the players who scored less runs but

took many wickets. So here we interpret

these two clusters as batsmen and

bowlers. The important point to note

here is that there were no labels of

batsmen and bowlers. Hence the learning

with unlabeled data is unsupervised

learning. So we saw supervised learning

where the data was labeled and the

unsupervised learning where the data was

unlabeled. And then there is

reinforcement learning which is a

reward-based learning or we can say that

it works on the principle of feedback.

Here let's say you provide the system

with an image of a dog and ask it to

identify it. The system identifies it as

a cat. So you give a negative feedback

to the machine saying that it's a dog's

image. The machine will learn from the

feedback and finally if it comes across

any other image of a dog, it'll be able

to classify it correctly. That is

reinforcement learning. To generalize

machine learning model, let's see a

flowchart. Input is given to a machine

learning model which then gives the

output according to the algorithm

applied. If it's right, we take the

output as our final result. Else we

provide feedback to the training model

and ask it to predict until it learns. I

hope you've understood supervised and

unsupervised learning. So let's have a

quick quiz. You have to determine

whether the given scenarios uses

supervised or unsupervised learning.

Simple, right? Scenario one. Facebook

recognizes your friend in a picture from

an album of tagged photographs.

Scenario two, Netflix recommends new

movies based on someone's past movie

choices.

Scenario three, analyzing bank data for

suspicious transactions and flagging the

fraud transactions. Think wisely and

comment below your answers. Moving on,

don't you sometimes wonder how is

machine learning possible in today's

era? Well, that's because today we have

humongous data available. Everybody's

online either making a transaction or

just surfing the internet and that's

generating a huge amount of data every

minute and that data my friend is the

key to analysis. Also, the memory

handling capabilities of computers have

largely increased which helps them to

process such huge amount of data at hand

without any delay. And yes, computers

now have great computational powers. So

there are a lot of applications of

machine learning out there. To name a

few, machine learning is used in

healthcare where diagnostics are

predicted for doctor's review. The

sentiment analysis that the tech giants

are doing on social media is another

interesting application of machine

learning. Fraud detection in the finance

sector and also to predict customer

churn in the e-commerce sector. While

booking a cab, you must have encountered

search pricing often where it says the

fair of your trip has been updated.

Continue booking. Yes, please. I'm

getting late for office. Well, that's an

interesting machine learning model which

is used by global taxi giant Uber and

others where they have differential

pricing in real time based on demand,

the number of cars available, bad

weather, rush hour, etc. So they use the

search pricing model to ensure that

those who need a cab can get one. Also,

it uses predictive modeling to predict

where the demand will be high with a

goal that drivers can take care of the

demand and search pricing can be

minimized. Great. Hey Siri, can you

remind me to book a cab at 6 p.m. today?

>> Okay, I'll remind you.

>> Thanks.

>> No problem.

>> Artificial intelligence, machine

learning, and deep learning represent

the evolution of computer science

towards creating intelligent systems. AI

is the broader concept striving to build

machines capable of humanlike

intelligence. ML is a subset of AI

emphasizing algorithms that learn from

data to make predictions or decisions.

DL in turn is a specialized branch of ML

that employs deep neural networks to

model complex patterns. Imagine an AI

powered voice assistant like Apple Siri.

It utilizes ML to understand and respond

to user queries, learning from

interactions over time. Deep learning

comes into play when Siri recognizes

speech patterns or interprets natural

language using neural networks to

process intricate features. The better

it becomes at understanding diverse

accents or refining responses

exemplifying the continuous learning

inherent in these technologies. AI seeks

to emulate human intelligence. ML

harness data for learning and DL employs

deep neural networks for intricate task.

The integration of these technologies

manifest in everyday applications,

transforming how we interact with and

benefit from intelligent systems. This

technology enables voice interaction,

allowing the device to play music, set

alarms, present audio books, and provide

up-to-date information on topics like

news, weather, sports, and traffic

reports, etc. Let's move forward and see

what is machine learning. Machine

learning is a subset of artificial

intelligence that focuses on developing

algorithms and models capable of

learning and making predictions or

decisions without being explicitly

programmed. ML systems leverage data to

recognize patterns, adapt and improve

their performance over time. There are

several types of machine learning.

Number one, supervised learning. The

algorithm is trained on a label data set

where each input is associated with a

corresponding output. Number two comes

as unsupervised learning. Unsupervised

learning deals with unlabelled data to

find inherent patterns or structures

within the information. And then comes

the reinforcement learning. This type

involves training agents to make

sequences of decisions by interacting

with an environment. And then comes

semi-supervised learning.

Semi-supervised learning combines

supervised and unsupervised learning

elements typically using a small amount

of labelled data and a larger pool of

unlabelled data. Let us move forward and

see what deep learning is. Deep

learning, a branch of machine learning,

focuses on algorithms inspired by the

human brain structure and functionality.

It excels in processing vast amounts of

both structured and unstructured data.

At the heart of deep learning are

artificial neural networks, empowering

machines to make decisions. The key

distinction between deep learning and

machine learning lies in data

presentation. Machine learning

algorithms typically demand structured

data while deep learning networks

operate through multiple layers of

artificial neural networks allowing them

to handle diverse data formats. So let's

start with the difference between

artificial intelligence, machine

learning and deep learning. And this

we'll show in a table form. So starting

with the definition.

So definition of artificial

intelligence. So broad field of machine

learning or creating machines with

intelligent behavior is artificial

intelligence. And when we talk about

machine learning, it's the subset of AI

focusing on algorithms learning from

data. And then comes the deep learning

that is specialized subset of ML using

deep neural networks. And now we'll see

the difference with the learning

approach between all these three. So in

learning approach artificial

intelligence can include rulebased

systems, expert system and more. And in

machine learning, it learns from data

patterns without explicit programming.

And then comes the deep learning where

it learns hierarchical representation

using neural networks. And if we talk

about scope, it encompasses various

techniques beyond learning from data.

And in machine learning, it primarily

focus on learning patterns from data.

And then the deep learning, it

specifically utilizes deep neural

networks for complex task. And now we'll

move to the next difference. And we'll

start with an example. So in artificial

intelligence, the example is autonomous

vehicles, chatboards or expert systems.

And for machine learning, it's spam

filters, recommendation systems, image

recognition. And in deep learning it is

image and speech recognition natural

language processing. And now we'll see

the difference for the data

requirements. So it depends on the

specific application and problem solving

approach. And in machine learning it

requires labeled or unlabelled data or

training. And for the deep learning it

relies on large amounts of labelled data

for training deep networks. And now for

the complexity artificial intelligence

addresses a wide range of task including

those beyond ML. And in machine

learning, it deals with moderate to

complex task depending on algorithms.

And for the deep learning, it is well

suited for intricate task often

requiring substantial computational

resources. And now see the flexibility.

So for the artificial intelligence, it

can be rule- based, evolving and

adaptive. And for the machine learning,

the flexibility adapts to patterns and

the changes in data. And for the deep

learning, it adapts to hierarchical

representations and diverse data types.

And then comes the training process. So

in artificial intelligence, training

process varies based on specific AI

techniques used. And in machine

learning, training involves feeding data

and adjusting model parameters. And in

deep learning, training involves

optimizing neural weights and

structures. And now we'll talk about the

applications between all these three

terms that is a IML and deep learning.

So for artificial intelligence the

applications are robotics, natural

language processing, game playing and

for machine learning it's predictive

analytics, fraud detection and

healthcare diagnostic and for the deep

learning that is image recognition,

speech synthesis and language

translation.

>> Now you guys must be thinking why should

I consider a career in AI? Well AI is

not just a passing trend. It's a seismic

shift that is reshaping our world and

creating new venues for innovation and

discovery. Now by embracing a career in

AI, you become a part of dynamic field

that thrives on solving complex problem,

pushing boundaries and making a profound

impact on society. The demand for AI

professionals is skyrocketing across the

industries from healthcare, finance,

entertainment, transportation.

Organizations are actively seeking

talented individuals who can harness the

power of AI and drive their business

forward. But what skills does it take to

become an AI engineer? How can you

embark on this thrilling journey? We

have the answer to all your questions.

Some steps are crucial to master the

field of AI and become an AI engineer.

Let's go through them real quick. So the

first step is to establish a strong

foundation in mathematics and

programming. Start by gaining a solid

understanding of critical mathematical

concept such as linear algebra, calculus

and probability theory. Additionally, it

is crucial to become proficient in

programming languages like Python which

is commonly used in AI and develop

coding skills. Next, you need to pursue

a degree in relevant field. Earn

bachelor's or master's degree in

computer science, data science, AI or a

related discipline to acquire a

comprehensive understanding of AI

principle and techniques and after that

you need to acquire knowledge in machine

learning and deep learning. Familiarize

yourself with ML algorithms, neural

network and deep learning frameworks

like for example TensorFlow, PyTorch to

train and optimize models using real

world data sets and afterward engage in

practical projects. Gain hands-on

experience and demonstrate your skills

by working on AI projects. Building a

portfolio of projects that showcase your

ability to solve AI problems can make a

strong impression on potential

employers. After that, collaborate and

network. This is really important.

Engage with AR communities, attend

conferences, and participate in online

forums to connect with professionals in

this field. Collaborating with others

can enhance your learning experience and

open up new opportunities.

Seek internships or entrylevel positions

where you can gain practical experience

through AI internships or entry-level

roles in industry or research

institution. Now this will provide

valuable exposure and help you further

develop your skills. After that

continuously learn and adapt. In the

fast-paced world of AR, it is very

important to stay updated on new

developments, explore specialized areas,

and embrace emerging technologies and

tools. Continual learning and

adaptability are essential for pursuing

a successful career as an AI engineer.

Now that you're familiar with the steps

involved in the journey of an AI

engineer, let's discuss the essential

skills you need to know to become an AI

engineer. So, here's a breakdown of the

skills needed. First one is having

strong programming abilities. This

typically refers to expertise in one or

more programming languages commonly used

in data science and machine learning

such as Python or R language. Now,

proficiency in programming allows you to

write efficient and scalable code for

data analysis, modeling and algorithm

implementation.

Next, you need knowledge of machine

learning algorithms. This involves

understanding and familiarity with wide

range of machine learning algorithms

including both supervised and

unsupervised techniques. You should be

able to select and apply appropriate

algorithms for specific problems as well

as evaluate and optimize their

performance. Next skill is proficiency

in statistics and mathematics. Sound

knowledge of statistics and mathematics

is fundamental for data analysis and

machine learning. You should be

comfortable with statistical concepts,

hypothesis testing, regression analysis,

probability theory, linear algebra and

calculus.

Now after that you have acquired a good

amount of knowledge of these skill set,

we'll move on to our next skill which is

having familiarity with deep learning

frameworks. Now deep learning has gained

significant popularity in recent years

and familiarity with deep learning

frameworks like TensorFlow, PyTorch or

Keras is valuable. Now these frameworks

provide tools and libraries for

building, training and deploying deep

neural networks for tasks such as image

recognition, natural language processing

and time series analysis. Next, you need

experience with big data technologies.

Dealing with large scale data sets

requires knowledge of big data

technologies such as Apache, Hadoop,

Spark or distributed computing

frameworks. Understanding how to

process, store and analyze data

efficiently in distributed environments

is very essential. Now after you have

gotten experience with big data

technologies, now it's the time to move

on to our next skill which is having

excellent problem solving and analytical

skills. Now these skills will enable you

to break down complex problems, identify

key factors and develop efficient

solution.

Now you should be able to adapt at

critical thinking, troubleshooting and

debugging to handle real world

challenges in data science and machine

learning. So guys, remember to stay

updated with the latest advancements in

the field and continue learning to stay

at the forefront of data science and

machine learning. So that's all we had

for you in this AI engineer road map. Do

you know how AI has become so fast? It's

now replacing entire teams in some

industries. Yes, it's true. Over 50% of

companies are already using AI to

automate jobs. AI tools are writing

emails, creating content, and even

giving job interviews. And while some

people are worried AI will take the job,

I let you in on a secret. AI is also

creating tons of highpaying roles. The

catch, you need the right skills to get

it. And that starts with learning with

the right programming language. Now,

I've tested a whole bunch of them.

Python, C++, R, Java, you name it. And

in this video, I'm breaking down the top

five programming languages for AI that

you need to know if you want to build a

career, land real jobs, and actually

stay relevant in the age of AI. We will

cover what each language is best at, how

to start learning, what kinds of AI jobs

they lead to, and yes, how much you can

earn with each one. All right, first up,

we've got Python. And honestly, this one

is the most valuable player of the AI

development. Just like the star player

in a sports team, Python is the go-to

language that everyone relies on when it

comes to building AI system. So, why is

Python the AI king? Let me break it

down. Simplicity and readability. Now,

Python is super easy to learn. It's

almost like writing in plain English.

You don't have to worry about

complicated code. If you're just

starting out in programming, that is

definitely the language you are going to

feel most comfortable with. It's got

this userfriendly vibe that makes it

simple even for people new to coding.

Second of all, it has got endless

libraries. Now, Python is packed with

tools. We call it libraries like

TensorFlow, PyTorch and Scikitlearn.

Think of these library as pre-made

toolkits that make AI development way

easier. They save you a lot of time

because instead of building everything

from scratch, you can use these

libraries to quickly train your models

and run algorithms. It's like having a

shortcut to building AI system. It has

also got rapid prototyping. If you need

to test your ideas quickly, Python is

perfect for that. You can build a model,

a simple version of your AI system in no

time. So whether you're working on

machine learning models or neural

networks, fancy word for AI system that

learn like the brain, Python let you

prototype or build a quick model fast.

So I know you must be wondering now what

kind of AI jobs can Python land me? It's

a great question. With Python, you could

land jobs like data scientist, machine

learning engineer, or an AI researcher.

Now these jobs typically pay between

around six lakh to 15 lakh peranom.

That's the salary range. But the best

part is as you gain more experience and

expertise that number will go way

higher. So how do you start learning

Python? You don't have to break the bank

to learn Python. You can get started

with free platforms and YouTube

channels. Simply learn even offers a

free comprehensive course in Python and

I'll leave the link for you to check it

out. And the best part is Python has got

huge community. So if you ever feel

stuck, there's always someone out there

who's ready to help you. Next, we'll

talk about C++. Now C++ isn't as

beginner friendly as Python, but it's a

beast when it comes to performance heavy

applications. If you're working on

realtime AI like self-driving cars or

high frequency trading algorithms, then

C++ is where you want to be. But why did

we choose C++ for AI? First of all,

because of its speed and efficiency.

Now, C++ is all about its speed. It's

the language you want when you're

working with large data sets or AI

applications that need to be super fast.

Second of all, it has got lowlevel

memory management. Now, C++ gives you

full control over memory, which is

essential when you're building AI system

that require extensive computation and

realtime performance. But isn't C++ more

complex than Python? Definitely, yes.

But if you're diving into AI

applications that require high

performance, think computer vision or

robotics, C++ is unmatched. It's a bit

trickier to learn, but if you want to

build realtime AI systems, it's worth

the effort. Roles like AI software

developer or computer vision engineer

are your goto with C++. The salary range

for these roles is around 8 lakh to 20

lakh peranom depending on the project's

complexity and your experience. Third on

a list is Java. This one's a workhorse

in the world of AI. And if you're aiming

to work on enterprise level AI projects,

then Java is definitely a language you

want to know. Now, it's not the first

choice for small scale AI projects. But

when it comes to big scalable systems,

Java is untouchable. So why Java for AI?

Because of its scalability. Now, you can

think Java as a beast when it comes to

handling large scale applications. If

you're working on AI system that need to

process huge data sets or manage complex

computations, then Java can handle it

all without breaking a sweat. It's

designed to scale which makes it perfect

for enterprise level AI projects where

big data is involved. Platform

independence. One of the best things

about Java is its right ones run

anywhere feature. It doesn't matter

which platform you're using, whether

it's Windows, Mac, Linux, Java can run

all of it without issue. This is a huge

win when you're building AI systems that

need to operate across multiple

platforms. Mature libraries. Java has

been around for decades and because of

that, it's packed with reliable

libraries for AI. Libraries like Qua,

H2O make implementing machine learning

models or building AI system a lot

smoother. These libraries we tried and

tested so you know you're working with

solid tools. Let's talk about what jobs

can you actually land with Java. Now

with Java you're looking at some big

roles in the AI world. Think of AI

solution architect or AI backend

developer. These positions are not just

highly respected but it also comes with

a solid salary range typically between 7

lakh to 18 lakh peranom. And with

experience, well, let's just say that

number can easily climb higher. Now, you

must be thinking, how do I get started

with Java? Now, if you're already

familiar with object- oriented

programming, learning Java will be a

breeze. And if you're new to it, don't

worry. You can start with some great

resources like a YouTube channel or

LinkedIn Learning. There are plenty of

courses that will teach you how to use

Java for AI from the ground up. Now,

let's talk about R. This one's for all

data science enthusiasts out there. If

you're diving into statistical AI and

the language built specifically for

handling massive data and performing

complex statistical analysis, R is your

goto. So why R for AI? Because of its

statistical power. R is packed with

tools for statistical modeling. So if

you're working on AI projects that need

data analysis before you even start

applying machine learning, R makes it a

breeze. It's got everything you need for

analyzing trends, finding patterns and

building strong predictive models. It

has also got a feature of its data

exploration and visualization. One of

the R's strength is its data exploration

and visualization capabilities. You can

easily plot, chart and analyze your data

to uncover insights. This makes art

perfect for the datadriven side of AI

development where understanding your

data is just as important as building

the models. But wait, can I still work

in AI if I learn R or is it just for

data analysis? Absolutely. R is

fantastic for AI projects that rely on

statistical methods and data analysis.

It's actually the language of choice for

roles like AI data analyst or

quantitative analyst where you'll be

building predictive models or analyzing

data trends to make decisions. Now these

roles are in high demand and the salary

range typically falls between 6 lakh to

12 lakh peranom but with experience you

can definitely push those numbers

higher. Now to get started with art, you

can find tons of free resources on a

plat new kid on the block that's growing

fast in the AI space. It's relatively

young compared to Python or C++. But

trust me, it's making a huge impact. And

here's why. Now, Julia was created back

in 2012 by a group of researchers who

wanted a programming language that could

handle the complex calculations required

for scientific computing and they nailed

it. But why did Julia grew so fast?

Well, it's been picking up speed because

it combines the performance of C++ with

the readability of Python. You get the

speed and efficiency that C++ is known

for, but with Python's clean and easy to

write code, it's like the best of both

worlds. Let's talk about why did we

choose Julia for AI? Because of its

speed and simplicity. Now, Julia's speed

is one of the biggest advantages. It's

designed for high performance computing.

So if you need to run complex AI models

or process tons of data, Julia will do

it in a fraction of the time it would

take in other languages. And the syntax,

it's also super easy to read and write.

So you're not sacrificing convenience

for performance. It has also got the

feature of scientific computing. Now

Julia is optimized for AI task like deep

learning and numerical analysis. You can

think AI applications in robotics, data

science, and scientific research. Now,

if you're working on projects that

require heavy computations or advanced

AI models, Julia's is your go-to. So,

why isn't everyone using Julia yet? It's

still growing, but Julia community

expanding rapidly, and more libraries

and frameworks are being developed every

day. It's quickly becoming a top choice

for high performance AI, and it

continues to evolve. And of course, I

expect to be even more popular. So, is

Julia better than Python or C++? Now the

answer is it depends. Now if you're

building scientific AI applications that

require high performance, Julia is a

fantastic option. It's still growing but

the community expands. Julia will

quickly become more powerful in the AI

space. Julia is perfect for roles like

AI developer in the scientific or

numerical computing space. Salaries can

range from 7 lakh to 15 lakh peranom

especially if you're working with

advanced AI. So guys there you have it

the five best programming languages for

AI. So whether you're interested in

machine learning, realtime AI or data

science, there's language for you. Each

of these will help you land AI job you

want and give you the tools you need to

build powerful AI system. Which one are

you going to start with? Drop your

thoughts in the comment section below

and let's talk about it. And if you

found this video helpful, hit that like,

share, and subscribe button to get more

AI tips and career advice by simply

learn. get started with the onboarding

and interface including the subscription

plan. As you can see here, it is

offering us three major plans. Now, now

there is a free version of manuals you

can use on a day-to-day basis. But make

sure to know that everyday credits are

six rupees.

>> But make sure everyday credits are only

300 to limited. But only 300 credits

will be assigned to you on a everyday

basis. Now, the first plan is $20 per

month, which gives you 300 fresh credits

every day, 4,000 credits per month,

in-depth research for everyday task,

professional website for standard

outboard, insightful slides for regular

content, task scaring, and wide

research, early access beta features,

and 20 concurrent task, 20 schedule

tasks. Now again if you are working in

an organization which where you can auto

you have to auto too many stuffs you can

upgrade to a $40 or $200 plan. Now $20

is for a person single usage because

it's only 300 fresh credits per day.

It's total of 4,000 per month as well.

So when it so when it comes to $40 plan

you can consider sharing it with two to

three people. Again it's 300 credits but

8,000 credits per month. all the other

things plus plus you'll get an addition

of 4,000 more credits to work on. Now

when it comes to 200 you'll get a

firstly you'll get free cloud computing

where you don't have to worry about the

storage and stuff and here it is 40,000

credits per month an organization which

uses automation tools a lot more can use

this now we are going to start by

understanding the manusi interface and

the first thing we need to lock out at

the hub which is basically your main

dashboard. Now before we start giving

task to manus AI it is very important to

understand how credits works because for

many users credits can be a little

confusing in the beginning. When you

open the dashboard you will notice that

man's AI shows two different credit

counters. The first one is the daily

refresh credits. These are the credits

that refresh every day. For example you

may see around 300 credits per day. The

important thing to remember is that

these are use them or lose them credits.

That means they reset every 24 hours and

if you don't use them, they do not carry

forward in the next day. So these are

the daily credits which are good for

regular task, quick experiments, small

research work, testing prompt or even

trying out different features inside

Manusa. The second credit counter is

your monthly pool. This is your main

credit balance for the month. For

example, if you're on a standard plan,

you may need something like 4,000

monthly credits. These credits are more

useful for larger and more complex task.

So, if you ask manus AI to do something

longunning like researching a topic

deeply, creating a report, browsing

multiple sources, analyzing information,

or even completing a multi-step

workflow, then this monthly pool gives

you the main runway to complete those

bigger tasks. So just remember this

simple difference. Daily credits are for

everyday use and reset every 24 hours.

Monthly credits are your larger credit

pool for bigger tasks throughout this

month. Now the next important thing in

the dashboard is the active task window.

This is where manusci shows the tasks

that are currently running and this is

one of the most powerful parts of the

platform.

Unlike a normal chatbot where you can

ask one question wait for one answer,

Manos AI can work on multiple task at

the same time. For example, on this

subscription you can run up to 20

concurrent task at once. That means

manos can work on multiple request in

parallel. Maybe one task is researching

a topic, another is preparing a

document, another is analyzing a website

and another is organizing the

information. For the free users, the

limit is usually lower around five

concurrent tasks. But the important

thing is not just the number of tasks.

The important thing is that these tasks

are asynchronous and cloud-based. This

means once you start a task, Manus AI

continuously working in cloud. You do

not have to keep watching the screen the

entire time. You can start a task, close

the browser, disconnect from the

internet, and even come back later. and

Manus AI can still continue to work on

that task in the background. This is

what makes it feel less like a normal AI

chat port and more like an AI worker.

You're not just asking a question and

waiting for the reply. You're assigning

work, letting the agent process it, and

then checking the results once the task

is completed. So before using Minus AI

for real workflows, always understand

these three things. Your daily credits

reset every day. Your monthly credits

support bigger and longer tasks and your

concurrent task window shows how many

jobs Manus AI is currently handling for

you. Once you understand this dashboard,

it becomes much easier to manage your

credits, plan your task properly, and

use Manus AI more efficiently. Now that

we have understood the dashboard and

credits, let's move on to the next

important part of Manus AI interface,

which is the goal, input, and task

planning area. Now this is where you

actually start working with manus. In a

normal chatbot we usually give a small

instructions one by one. But in Manus AI

the idea is slightly different. Here you

give a highle goal and manus plans the

steps needed to complete that goal. So

in the main input box let's type a

simple goal such as research the top AI

tools for content creation and create a

comparison report. So let's start.

research the top AI tools for content

creation and create comparison report.

Now here we have assigned a proper goal

to Manus AI. Now notice what happens

after we enter this prompt. Manos does

not directly jump into the final answer.

First it create a task plan. This is

where you will see a to-do list or

step-by-step structure showing how Manus

is planning to complete the task. For

example, manus may break the goal into

steps like understanding the topic,

searching the AI content, creation

tools, collecting useful information and

comparing those tools and finally

preparing the report. So here you can

see the steps. In simple words, manus is

taking one big goal and breaking it into

smaller actions. This view is very

important because it gives us a chance

to review the plan before the agent

starts doing heavy work. Before manos

begins browsing, opening pages,

analyzing sources and consuming more

credits, we can quickly check whether

the plan looks correct. For example, in

this case, we should check is manners

searching for the right type of tooth.

Is it planning to compare them properly?

Is it going to create a final report as

we asked? If the plan looks correct, we

can continue. But if the plan looks

incomplete or slightly wrong, we can

stop and adjust the prompt before moving

forward. This helps us avoid wasting

time and credits. So the key point here

is simple. In manus AI, we don't need to

write every step manually. We can give

one clear goal and manus will create a

plan for completing it. But before

allowing the task to continue, always

review the documentation decomposition.

But before allowing the task to

continue, always review the

decomposition view. This helps you

understand how the agent is thinking and

whether it is moving in the right

direction. So in this example, our goal

was to research the top AI tools for

content creation and create a comparison

report. And Manus turns that single bowl

into the structured task plan that can

review before execution. This is what

makes Manus air different from a regular

chatbot. It does not just answer

immediately. It plans the work first,

shows the direction and then start

completing the task. Now that Manus has

understood our goal, the created task

plan, the next step is execution. It's

already executing. This is where Manus

AI actually starts working on a task.

You can think of this part as a hand in

the platform. The goal input is where

Manus understands what we want. The

planning view is where it decide how to

do it and the execution view is where it

actually performs the work. Once we

approve or continue with the task,

manage begins completing the steps one

by one. The interesting part is that we

can watch this happen in real time. On

one side, you will usually see the

progress list or task steps. This shows

that manus has completed what is

currently doing and what is still

remaining. The next is that you can see

manus actually taking action. For

example, if the task is repeat, for

example, if the task requires research,

you can see the agent opening websites

and browsing pages. If the task requires

collecting information, it may take

screenshots, extract details, or even

organize the data. If the task needs a

structured output, manuals may update a

spreadsheet, write content, run the

code, or even build an interactive

artifact. So instead of only showing the

final result, manos shows the workflow

while it's happening. This is useful

because we can understand how the agent

is working not just what answer it gives

to the end. Now another important thing

is to understand here is the sandbox

environment. Manos does not directly

operate your local computer. It works

inside a cloud and is created for the

task. Inside this sandbox, miners can

browse websites, collect information,

test the ideas, run code, fill forms and

build outputs without affecting your

personal system. For example, if we ask

miners to research AI tools and prepare

a vision report, it can browse different

website, collect the required details,

organize them and then create the final

report inside this workspace. And for

more advanced task, the sandbox can also

help maners create things like websites,

slide decks, spreadsheets, dashboards,

and other interactive files. This is one

of the major reasons maners feels

different from the normal chatbot. A

regular chatbot mostly gives a text

responses. But maners can actually

perform actions inside a controlled

environment. So while the task is

running, we should keep an eye on two

things. First the progress list to

understand which step manus is working

on. Second is realtime action view to

see what the agent is actually doing.

This makes the whole process more

transparent. You're not blindly waiting

for the final output. You can see the

agent browsing, checking information,

organizing the data and building results

step by step. So in simple terms, the

execution view shows manus in action.

The sidebyside workflow helps us track

the task in real time and the sandbox

environment gives Manus a safe cloud

workspace where it can browse, run code,

collect data and create useful outputs.

This is the part where Manus moves from

planning the work to actually doing the

work. So we'll get back to this task

once it is completed. Let's start with a

new task. Now that we have seen how

Manus works inside the browser, let's

look at how can Manus be on normal web

interface. Now here you can even connect

a different apps such as Gmail, browser,

meta and you can add other connectors as

well if you're planning to automate any

kind of workflow. Now here when you come

to the desktop side you will have a

mobile app as well as the desktop app as

well. Now if you come to settings you

may find an option called integration.

This is where you can actually connect

manos with platforms like slack,

telegram or even line. So as you can see

here there are connectors. This is

useful because it allows you to interact

with manus through a messaging apps you

already use. For example, instead of

opening the browser every single time,

you can just delegate a task, check the

progress or monitor updates from a

messaging app. So if you're working with

a team, Slack can be useful. If you want

quick mobile access, WhatsApp, Telegram

or Lion can make it easier to stay

connected with the agent. Part two manus

AI. Let's continue. The main benefit is

remote control. You can start monitoring

task even when there is no sitting in

front of the main system. Now the next

advanced feature is the desktop my

computer feature. You can download the

computer version here in the desktop

app. This is available when you have

Manus desktop app installed in your Mac

or a PC. Here Manus can request access

for your local machine for specific

action. For example, it may need you to

read a local file, open a folder or run

a terminal command. But the important

thing is to notice that manus does not

get a fully access automatically. There

are permissions grades. There is a

permission gate when manus wants to

perform an action on your computer. You

will see prompts like allow once or

allow always. From a safety point of

view, allow once means you are giving

permission only for that specific

action. Always allow means you are

allowing that type of action more

regularly depending on the setup. So

while showing this, this is especially

useful when you want manos to work with

files on systems, run scripts or even

help with local development tasks. Now

the third advanced area is the web app

builder. This is where manage becomes

even more powerful. In the web app

builder interface, you can see manage

generating a live interactive web

application. This is not just writing a

text or giving code snippets. It can

actually build pages, connect the

databases, structure the app and prepare

it while working with the project. For

example, if we ask manus to create a

simple landing page or a small web app,

it can generate a layout and add

interactive sections, connect the

required backend logic, and even support

things like database setup and SEO

optimization. The best part here is that

you can watch the agent work step by

step. You can see it creating files,

updating the design, testing the pages

and even improving the final output. So

this part is useful for users who want

to build something practical like

websites, dashboard, internal tool,

product page or even prototype without

manually writing everything line of code

from scratch. To summarize this section,

so now let's test the same logic. Now

let's ask minus AI to create a web

landing page for a skincare brand. So

create a brand. Now to summarize this

section, the browser is the main place

where you use manus AI which is this.

The messaging integration help you

delegate and monitor task remotely. The

desktop app gives you manus control

access to your local machine with

permission prompts. And the web app

builder helps manus create live

interactive web project. So this is what

takes manus from being just a web- based

AI agent to something that can connect

with your communication tool, your

computer and your real project works.

Now as you can see there are approaches

here. This is the code for the entire

web page. Let it generate. I'll show you

the output since this is just running in

the first step. There are more three

steps involved in this. So we'll get

back to this once this is done. Now we

are going to see where Manus AI becomes

really powerful which is deep research

and data. The main idea here is very

simple. Manus AI is not just a chatboard

that gives one quick answer. It can work

more like an autonomous research worker.

That means you can give it a goal and it

can plan the task, browse multiple

sources, collect the information, cross

the check details and organize the

findings and finally create a proper

output. So instead of manually opening

20 tabs and copying the nodes, checking

the resources and building the report

yourself can handle a larger part of

that workflow for you. Let's start with

a we'll just use a practical prompting

as of now. Now for this demo, you can

just type in research the top CRM tools

for small business and create a

comparison report with pricing, key

features, pros, cons, best use cases and

source link. So I've given the exact

same prompting. Now once we enter this

goal, manus first creates a plan. This

is important because the task is not

just asking for a simple answer. We are

asking manage to research multiple CRM

tools, compare them and prepare a

structured report. Once the task starts,

notice how manners does not depend only

on one search result. It begins visiting

different websites and checking product

pages, pricing pages, review platform,

blogs, and others available sources.

This is what we call multi-source

research. For example, if MinusAI is

researching CRM tools, it may check

official websites for pricing, review

platforms for user feedback and

comparison articles for feature level

difference. The important thing here is

that Manos is not just collecting random

information. It is trying to cross

interface the details. So if one website

mentions a price, Manos can compare it

with the official pricing page. If one

source mentions a feature, it can check

whether the same feature is also listed

on the product website. This helps

improve the quality of the research.

Now, while the agent is working, keep

your attention on realtime interaction

view. On one side, you can see the task

progress. On the other side, you see

minus browsing websites, opening pages,

taking screenshots, reading the

information, and updating its findings.

This makes the process more transparent.

You're not blind. You're not blindly

waiting for final answer. So you are

actually seeing how the agent is

collecting and organizing the

information. Another important thing is

to notice how manage handles small

problems during the search. Sometimes a

page may not open. Sometimes a link may

be broken. Sometimes a website may be

JavaScript heavy and difficult to read.

In manual workflow we would have stopped

and find another source assets. But

maners has planning layer that can

create recovery steps. So if one source

does not work, it can try another

source. search again or adjust the path

without needing constant human help.

This is why manners is useful for

research heavy tasks. At the end, the

output should not just be a paragraph

summary. A good result should be a

structured artifact like a comparison

table or a full research report. For the

CRM example, the final output can

include tools, names, pricing, key

features, pros, cons, best use cases,

and source links, which we'll check back

in a few minutes. If you can move beyond

one short answers and prefer a full

research workflow across multiple

sources, manusi is the tool. Now let's

move on to which is wide search. This is

more advanced credit intensive feature.

So we'll get back to all the three in a

minute. We'll get back to all the three

outputs and I explain what was the exact

steps required. So coming back to wide

research. In normal research, the agent

may explore sources step by step. But in

wide research, the idea is very

different. Wide research is designed for

scaling. Instead of checking a few

sources one after the other, it can

explore many sources in parallel. Manus

described wide research as using

parallel multi-agent orchestration where

many agents can work across large

research space at the same time. So this

is not meant for basic questions like

what is CRM or even give me five tools.

This feature is better for high impact

research tasks like market analysis,

competitive research, industry reports,

investment research, product research,

or even strategy planning. For example,

we can use a large version of the same

CRM topic. Run wide research on the CRM

software market for smaller businesses.

Compare major players, pricing, trends,

AI features, and even customer

sentiment, market positions, or even a

growth opportunities. This kind of

prompt is much broader. Here we're not

only asking for a tool comparison. We

are asking miners to understand the

market from a different angles. It may

explore companies, websites, review

sites, market reports, competitors,

pages, product documentation, user

discussions, and other public sources.

Now, before starting wide research,

always explain the credit part clearly.

This type of task can consume a lot more

credits than a normal research would.

Since wide research explores a large

number of sources and runs a much

heavier workflow, it can cost

significantly more credits. So we should

use it for important research work, not

for a simple Q&A. This is important for

learners. Think of it like hiring a full

research team for one task. You would

not only use them for a small

definition. You would use it when the

output has real business value. So the

main takeaway here is use normal

research for focused task. Use why

research when you need a large scale

high depth analysis across many sources.

Now next we'll move on to it is useful

because it shows how manuals can move

from raw data to a finished business

report. Here for example let's say let's

upload a CSV file. So here I've taken a

random data set from Kaggle and I've

uploaded it. It says loan data set. Now

let's give it a prompt saying analyze

this loan data and create a report

showing revenue trends, top performing

products etc. So let's just say analyze

this loan data and create a report

showing the trends. So mind you I have

already cleaned this data and executed

using AI which is in collab but still it

took me like proper an hour to create

it. So let's just leave it. Now as you

can see this is where maners becomes

different from normal AI tools. It does

not only look into the file and guess

the answer. It can work inside a cloud

sandbox. Inside this sandbox, miners can

write and execute code such as Python to

process the data. So if the file

contains thousands of rows, miners can

calculate totals, averages, trends,

category performance, product

performance, regional performance, and

other useful metrices. While this is

happening, show the executional view.

You may see man is reading the file,

writing the code, running analysis,

checking the output and generating

charts. This is manus is not producing

text. It is actually performing mini

data and this is workflow. After

processing the data, Manus can also

create visualization. For example, it

can generate charts showing loan

prediction data, which category will

take more loan, etc. Then the final

step, it can synthesize everything into

a business report. Now, what does a good

report include? what the data shows,

which products are performing well,

which areas need attention, which trends

are visible and what actions the

business should take next. So from one

uploading of file and one prompt, Manus

can complete an end to end workflow. It

can pass the data, run the code, create

charts, interpret the results and write

a final report. This is why Minus is

very useful for business users,

analytics, marketers, sales teams,

founders, and students learning data

analysis. Now let's see how Manis AI can

work on autonomous research worker. It

can browse multiple sources, analyze the

data, run code and prepare structured

reports. Now in this module we will see

manus AI as a creator. This is where

manus move from just giving answers to

creating finished functional artifacts.

So instead of only asking manus to

explain something, we can ask to build

something. It can create web apps, slide

text, posters, infographic, visual

content and also complete project assets

from a single natural language prompt.

So let's get started. So here let's give

manners a single prompt. Now let's ask

it to create a landing page for AI

productivity tools for students with

sections for features, pricing,

testimonials, FAQs, and call in action.

Now can you notice what happens here? We

are not giving miners a full design

document. We are not writing code. We

are not explaining very section step by

step. We're only giving it an idea.

Manus takes this idea, understands the

goal, creates a plan, decides the page

structure, writes the content, designs

the layout, and starts building the

page. This is important. Manus is not

just giving us text response. It is

creating a clickable portfolio. So, as

you can see, it already started creating

the This can be very useful for

developers, product managers, startup

founders, marketers, and business teams.

If someone has an idea and want to

quickly see how it might look at a

website, manners can help create the

first version very quickly. Instead of

spending hours preparing a wireframe or

explaining the idea to a designer or a

developer, we can just use maners to

create a rough working version. Then we

can share it with the team, client or

stakeholder for feedback. Depending on

the tunnels can also help with more

advanced parts like databases, payment

flow, SEO friendly structure and

deployment related steps. But for

beginners, the main thing is to

understand this manual can move from an

idea to a functional prototype. Also,

this type of task is more resource

inensive than a simple chat response.

Building a web page may consume hundreds

of credits. Sometimes around 500 to,000

or even more depending on the

complexity. So before running a web app

task, always check the estimated credit

usage. This is because minus is not only

writing text, it is planning, coding,

testing, building and sometimes handling

deployment steps as well. Now let's move

on to the next part. Now, now let's ask

manus to create a slide deck. Now let's

ask the manus AI to create a text on the

future of AI agents for business teams.

Now once we give this prompt, Manus

starts planning the slide tech. It does

not randomly create slide. It first

creates a proper structure. For example,

it may begin with an introduction, then

explain what AI agents are, why

businesses are using them, their

benefits, use cases, challenges, and

finally a conclusion. This is what makes

the output useful. It's not just a set

of separate slides. It's a structured

visual story for research heavy topics.

Miners can also browse the web, collect

useful information and include cited

points. This makes it useful for

business presentation, research decks,

pitch decks, training models, and

internal reports. While the task is

running, look at the interaction view.

You can see manners creating the

outline, preparing slide content,

improving the design, building the final

deck and once the deck is ready, you can

usually download in a businessfriendly

format like Pex. So we can still open it

in PowerPoint and make final manual

changes. Now this is very important

because manus gives us a strong first

version but we can still fine-tune it

the slides based on our brand audience

or even presentation style. Now let's

move on. Let's come back to this later.

Let's see what are the outputs for all

the prompting that we have given. So

firstly I have asked it to create a

landing page for a skincare brand. Now

as you can see there is a skincare brand

where you can also edit these. So the

name is given the benefits products

purifying tensor what is the cost in

dollars. You can edit the landing page.

This usually used to take days for an

UIUX designer to design the entire page.

is just done with a small prompt. So you

have products, reviews, shop now and if

you come here ready to transform your

skin, the shops, new arrivals, colle

collection, support, etc. This doesn't

look like it's just done from a

prompting. Now if you come to the second

one, let's we had asked to compare the

top CRM tools. Let's see what's the

answer for that. So here the prompt was

to research the top CRM tools for small

businesses and create a comparison

report with pricing, key features, pros,

cons, best use cases and course link. So

as you can see let's open this report.

So here we have a summary where small

business CRM section is the best

approach as a trade-off among these easy

tools. Now is it comparing all the

things that we have given? The first one

it's HubSpot sales hub starting price

key features pros cons best use cases

and source link all the things are

present usually if you use a person they

used to browse through every single

website they could find and create such

kind of report now it's done in just a

small prompt that I've given now let's

move on to the next one which is loan

data analysis this is the most useful

tool for data analyst because we spend

hours. They spend hours cleaning the

data, visualizing trends, what graphs is

suitable for what kind of data,

normalizing the data and so many other

steps. Now, if you can just upload a

file and ask it to create all the

reports and all the things necessary to

take a business decision, this will be

the most useful tool for data analysts.

Now, let's see the answer for this. As

you can see the graphs are there. Let me

just open. You can give a prompt where

which kind of graph you want, what

against what graph you want etc. You can

see the credit history, marital status,

property area, education,

self-employment and dependency all

against approval rate. Now if you come

here there is a summary as well which is

a report. Now the summary is that the

report analyzes 614 do applications

using the uploaded loan data set

covering applicants demographic income

co-licant income etc. The data set shows

422 applications were approved

presenting an overall 68% while 192

applications were rejected representing

a rejection rate of 31%. Now as you can

see we have approval and rejected rate

and the data overview. What are the

data?

Now here this is a very small data set

and I took almost an day to work with

this data set and create modeling etc.

This is done within a few minutes and

this is amazing because it takes a lot

of time cleaning the data set knowing

the data how to understand the data.

This sorts out all the problem. Now

coming to the next one content creation.

So here I had asked manusi to research

the top AI tools for content creation

and create a company report. So here as

you can see there is a report that is

given. Let's preview it. So here the

heading is there explore tools download

the report again this is treating as

like a website that has all the

information. So here you can see tool

distribution by category text generation

tool is like one etc. Pricing tier

distribution 63.2 to AI tools directly.

The first thing is chat GPT which is

probably mostly consumed I think. Next

is Jasper AI. Then we have Canva AI and

next Grammarly Ptory Morph AI Descript

Midjourney

Surfer SEO Gemini Claude Copy.ai etc. So

here you can see the price also what is

best suited for there is a free version

also. So it's given free version as well

rating. This is amazing for content

creation because usually we don't get

pictures which give the exact direction

or exact ratio of the exact numbers that

we found online. It's either we have to

create from scratch. So this is amazing

for content like you have ratings, you

have pricings, you have to compare them,

select the top tools to compare. Let's

compare chat chibity and Jasper sorry

Jasper and chat chibity and also Canva

AI all three are equally used. Now let's

deselect them and copy.ai. Now as you

can see copy.ai is a little bit less on

ratings. Oh my god this is too good to

be a tool. This is literally AI to work.

Let's check out the next one which is

landing page for AI productivity. Now

again this also will be a landing page.

So it's basically like a website. Now as

you can see we have the heading college

study flow features pricing testimonials

FAQs study flow AI is your personal AI

tutor study planner productive companion

get instead explanation organize your

listings there's a free trial watch a

demo powerful features for the success

everything you need to excel in your

studies all in one place etc. And you

have the pricing as well. This looks

like a legit platform website that has

no flaws. There is literally a review

rating also frequently asked questions

which is common in most of the websites.

And then you have the down at 2024 study

flow AI all rights reserved. Next let's

see if the slide deck is ready. Now

let's play the PPT. It's about the

future of AI agents for business. So as

you can see first is the heading

footages. The next one is core ship with

this AI agents change the unit of work.

And then you have what and all things

are changing. Why now the agent stack is

maturing. AI agents are not just smarter

chat bots. What are the difference

between chatbot co-pilot AI agent agent

portfolio? The new team model in human

agent collaboration business teams will

adopt agents by functions scale agents

required enterprise architecture

governance adoption. It's a legit PPT to

explain each and every single step of AI

agents future. Manus AI is literally

describing how AI is put to work not

just give a text response.

>> All right, so we're ready to start. uh

we are going to start with this you know

first course which is going to study the

basics of Python. Python will be our

primary focus for the entire program. Um

we will use co-pilot. So there's there

will be co-pilot material um later on in

the program but like in this first

course we're going to be focused on

Python and and for most um things we

will be using Python. Um even when we

use co-pilot it will produce Python code

everything we do will be in Python. I

think one of the things is by you know

by the end of the program if anything

else you guys will be in a much better

position with Python. You'll be better

Python coders by the end by the end of

the program. If you don't learn anything

else you'll get better at Python. I

promise. Uh because that's you know all

of our examples all of our demos

everything we do will be in Python. So

you'll you'll get better at it. uh for

sure and we'll have a lot of practice to

do that. Okay. So this first lesson is

all about an introduction to what Python

is. So if you're completely unfamiliar

with it, totally fine. We will uh get

you up to speed and talk about the

fundamentals and how to set everything

up on your own computer and talk about

the various ways to um utilize Python.

that some of it will involve a setup you

can do on your own computer. Some of it

will involve some cloud resources um so

that you don't need to set anything up

on your computer if you don't want to.

Um we'll have options there which will

be nice. So I will show us those and

walk us through those. But this first

lesson all about the basics uh and

getting set up. So, um what's

interesting is like at the beginning of

every lesson, we usually have this uh

kind of um engagement or discussion. Uh

but you know, we've I kind of already

asked you guys about this of uh uh if

you're familiar with programming, if

you're familiar with Python. Um but one

thing I want you to think about a little

bit is that um especially as we go along

and learn about what Python is is why is

Python the

chosen language for AI? So why is it the

one that everyone uses uh to do AI? And

I think what you're going to learn is

that it has a really amazing ecosystem

that has been around for a long time

that um supports AI in particular. So,

Python is the go-to for anything AI,

data science, machine learning, anything

in that sort. Uh, because it's been used

for so long for that and it has such a

uh community and ecosystem around it.

That's something we're going to learn.

It's also really easy to learn and use,

which makes it nice to to be uh kind of

an introduction to the field. It doesn't

take a lot to get started in it.

because it's so easy to work with. Um, I

can tell you as someone who's gone

through that experience, like I studied

mathematics in college and in graduate

school and studied like probability and

statistics, but I was able to teach

myself Python primarily and use that to

get into kind of data science and

machine learning in the industry.

So, and I think that's a common story is

people and I've seen that from many

learners coming from uh different

backgrounds. Uh they've been able to

pick up Python pretty easily because

it's a very easy language to understand

and and syntax of it and there's so many

tools within it that make it really easy

to work with.

So, um I promise it won't be as uh

daunting as it may seem even if you're

coming at it from zero experience. Uh, I

think you'll find this is the perfect

way to get into programming and get into

data science and and AI and machine

learning because it's so easy to pick up

and learn and it has such a nice rich

community ecosystem.

So, just wanted to mention that.

Okay. So, some of our objectives for

this first lesson will be to talk about

programming languages in general and um

programming in general. So maybe you

know more generic than Python just you

know what are what do general programs

look like? What are some of the building

blocks of programs that are important?

What are some of those uh key principles

of programming that we will want to

follow as well? Even if we're doing

Python for AI purposes.

Um so just talk about programming in

general and then kind of zoom in on

Python as we go along. One of the things

we'll be interested in doing is just

getting you guys set up. So talk about

how we can configure Python for you to

use on your own machine. Um but also

have some options that don't require

installing anything on your own machine.

Uh which is nice. Um and then as I said,

we'll kind of zoom in on Python, talk

about its benefits, uh some of the nice

features. I've kind of already mentioned

it. Really big community around it, easy

to learn. We'll just talk about those

more in detail. Talk about um why it's

so popular in the AI world. Um,

and then we'll get into some very

fundamental things specific to Python.

So once we talk about the background,

get you guys set up, we'll go into uh

some of the syntax basics, things like

identifiers, things like indentation,

comments, um, some of the basics of the

code that are going to be important for

you to kind of get started with. Um and

then talk about some of the basic data

types that Python offers to manipulate

and work with data which of course is

important um when you know as we go

forward and and do anything with data

which of course with AI we will be

interested in doing um but that's these

are the objectives of just the this

first lesson. As we go forward we're

going to learn about many other basic

topics within Python. So things like how

to write functions, how to build

objects, how to manipulate our flow of

the program with like things like if

else statements, things like loops.

We'll learn all about that in kind of

the next lessons after this one. But

this is all the content for this lesson.

I anticipate today

we will get through all of this today

and then get into the second lesson

which will um get into those kind of if

else and loops. So we'll we'll get we'll

I'm sure by today we'll get into those.

All right. Any questions on kind of what

we're going to learn in this first

lesson? So mainly trying to get you guys

set up, give you some background on

Python and then towards the end of the

lesson um get into some basics of the

syntax is kind of the goals I would say.

Okay. Okay. So when we talk about

programming um what do we mean by

programming in general? It's really uh

synonymous with instruction. So

programming really means giving or

writing instructions for a computer to

perform tasks. Um so these instructions

we write down in what we call code. But

those those are just telling the

computer what to do. And of course the

computer's not going to do anything

unless we write down these instructions.

So these instructions can do really

powerful things. They can power, you

know, whole applications, things that we

use every day like Microsoft Word,

PowerPoint, Excel, those kind of things.

Um they can automate tasks. They can um

power websites. Um they can do AI,

right? So we can have um things like

chat GBT and Alexa and Siri, etc., etc.

Um these are all powered by instructions

telling the computer what to do.

One of the things that we will get

better at as we go along is figuring out

how to write these instructions in

Python. Python is going to be the

language we write those instructions in

um and and they will be executed by a

Python um program. But we should think

of programming in general as just

instructing the computer what to do just

at a high level. Right?

So when we talk about these

instructions, they have two ways of

being executed by the the computer. Um

and roughly these break down into what

we call interpreted languages and

compiled languages. So that the code

that we write which is um representing

the instructions that we write can be

executed um in one of these two ways.

Let me start with the left. So the

interpreted languages.

This means that the computer is

literally executing the the instructions

line by line by line when we run the

program. So there is no

translation of anything. It's just

literally taking our instructions and

running it line by line, instruction by

instruction essentially. Um, now the

advantage to doing this is that it's uh

easier to debug because the instructions

are going to be executed one by one. So

it can hit an error pretty quick. If

there's a mistake in one instruction,

nothing else will run. Um, however, it's

also slower because we're going to take

it one instruction at a time. Um, and so

the the uh this way of running programs

tends to be slower, but it's also easier

to work with, which is why we're so

interested in Python. It's in this

bucket of what we call interpreted

languages. So a lot of scripting

languages find themselves in this bucket

of being executed one line at a time. No

translation needed by the machine. It

just reads our instructions and executes

it. The thing that does the execution is

called an interpreter.

Um, and Python has an interpreter that

we will get you guys set up with on your

own machine that can execute Python

code. So you need an interpreter. The

interpreter just executes your

instructions line by line by line. Um,

so some examples would be like Python.

That's what we're going to study in this

um entire program. But there's other

languages like JavaScript, Ruby,

um Pearl, many others that are uh

interpreted. They require an

interpreter, but they execute line by

line by line and there's no intermediate

translation of anything. Um it's kind of

executed as is. Now, contrast this with

compiled languages, which are uh kind of

a different piece. they these these

instructions have to be translated into

something the machine can understand in

order to execute. So there is an

intermediate step of what we call

compiling the code um into uh basically

a translated version of your

instructions so that the machine can

execute it. Now there's a trade-off

there. Doing that can make it more

difficult to develop and it can take

longer to debug because you have to go

through this translation step every

single time through the compiler.

But when you run the code because it's

already been translated into this

machine format, it's a lot faster. Um,

so some examples of languages like this

are C, C++, Java,

um, Go,

but uh, we won't really be working with

those. We'll just be sticking with

Python. But if you have experience with

those languages, those you're probably

familiar with this, you have to compile

the program first before you can execute

it. But we are going to be in this

interpreted world. If you know and it's

okay like if none of this makes sense,

that's okay. Just understand that um

generally interpreted languages are

going to be more user friendly because

they're they're easier to execute. They

don't require as many moving parts as

what a compiled language would require.

which is nice for us, right? Nice for

Python. That's what we're going to be

interested in working with. Uh kind of

um yeah, they're kind of rel So, so the

question is are JavaScript and Java

related? Kind of. Um, JavaScript is kind

of like the um the the

scripting version of um some of the same

concepts we see in Java, but Java is the

compiled um it it requires a a special

kind of what's called a Java runtime,

which is a a compiler to translate the

Java code into um machine code that the

Java runtime will execute. JavaScript is

not like that at all. It can actually be

ran in a web browser which is um

JavaScript usually powers a lot of like

front-end websites are usually powered

by JavaScript and Java usually powers

more like backend

um applications like actual software

programs are usually would be coded in

Java. JavaScript is going to be used

more for like building a website. But,

you know, I'm not an expert on that

really, but that's kind of my

understanding of it. And if anyone is an

expert on those differences, feel free

to let us know in the chat. But, uh,

that's my that's my basic summary of

that. Okay. So, we have interpreted

languages. That's where Python falls

under. So, it just um summarizing that,

it's going to be easier to work with

those, which is great for us. That's

another reason why Python's so easy.

It's interpreted, meaning that

everything executes. We don't need to

worry about compiling things, which is

nice. Um, but also in terms of

programming, there's also uh categories

of how the instructions are written that

you can bucket different languages into.

So for example um some language are are

more um procedural in nature meaning

that you write out all the instructions

exactly kind of line by line by line.

You don't really organize things at all

in your instructions.

Um so some examples would be like C and

Pascal

are more like that. Um then on the

opposite end of the spectrum is kind of

object-oriented

in which case you uh build your code and

organize it around the idea of

everything being an object. And so some

uh Python actually falls into this

category where um uh most things in

Python are objects and you manipulate

objects and objects have data to them.

They have things they can do and

interact with other objects. Um, so

think of it just as a way we will

organize our instructions.

Python allows us to organize it around

the concept of an object. We'll learn

about what that means as we go along,

but just realizing that some programming

languages break down along these um kind

of buckets here. Um, Python is also a

scripted language, meaning you can write

out your code in a individual script and

you can e that you can have an

interpreter that executes that script.

Um, so you don't need to organize all

your code inside of an object. So for

that reason, Python super flexible.

That's another reason why it's so nice

to use. It actually falls into both of

these buckets on the right, which is

very convenient. We can have basically

this means we can have a lot of

organization or very little organization

depending on how we want to set it up.

Yeah, Roberto. So even though there are

different types so Java is compiled and

Python is interpreted

um they are both object-oriented meaning

so think of the this slide as telling

you how the instructions are organized.

So how they are executed is different.

So, Java requires a a compiler to

execute things. Python requires an

interpreter.

This is more about how the instructions

are organized. So, Java and Python both

allow you to organize your code into

objects.

Um, but what's nice about Python is it

also falls under the bucket of

scripting, meaning that it allows you to

organize things into scripts, which is

less organization than it would be in

into objects. We're actually going to

learn about objects later on in a future

lesson, like how to build objects and

what they mean.

So yeah, even though they're different,

they're both object-oriented, which just

means that you can organize your code

into objects. Python allows that. So

does Java. So does C++.

Uh many many languages allow for um

organizing your your code into objects.

So we're going to learn about that.

It's It's not that one's better. They're

just um I I would put them at different

So, let me draw this. I would put them

at different spectrum, different ends of

the spectrum on organization.

So, scripting

is very loose. Basically, you it's more

like a an individual um uh set of

instructions to do one task. you can

just have and you can have many

individual scripts to do many small

tasks. Um, and then on the other end of

the spectrum, think about it as like

you've organized your cabinet into many

folders and many like uh you know many

pieces of organization that are we would

call objects. Um so objectoriented

programming OOP is kind of on the other

end of the spectrum when it comes to

like level level

of organization.

Does that make sense? So scripting very

loose. It usually scripting is is um

reserved for like one task and it's um

you're just writing out your

instructions to accomplish that one

task.

um which is helpful for like automation

of things because you're going you're

usually automating like a single task.

Um so it's very loose. It's not very

organized into nothing is organized

necessarily into objects. Um very loose

organization. Object-oriented is much

more structure to it and things being

put into objects um in order to

manipulate and work with objects

throughout the program. Yeah, it's not

that one's better. I think it's more

just use case dependent. Um there are

times where it actually will benefit us

from using objects. Um and I think the

thing to pay attention to on this slide

is that look at where Python falls into.

It actually falls into both. Meaning

that we can have things very loose and

easy to work with because scripting

usually will be faster and easier to

just write something to to accomplish

one task. But we have the flexibility to

organize our code into objects if we

want to. which will be better for

bigger tasks that require more

organization

like training a neural network or

building an LLM.

Those bigger tasks would benefit from

organization.

And then uh finally on this slide um

there are languages that are built on

the concept of um their their entire way

of writing instructions is more in a

functional way meaning everything is

based on operating uh functions and

variables. Um and so there are some

languages like that has and scholar are

very popular ones. Um but that is can be

very difficult to learn. It's it can be

difficult but very nice in some ways

because uh it can be very natural to

think of um you manipulate like giving

instructions to computer in a functional

way. Think about it as like applying a

function to a variable.

Um that makes sense but writing your all

of your instructions in that way can be

kind of difficult to learn. So for that

reason I think these languages are more

difficult to learn but they can be very

powerful. Um and they find themselves

very useful in like operating on big

data.

Um so if you ever heard of like Spark um

Spark operates with uh Scola for

instance um but uh we won't really focus

on functional. It's kind of its own

paradigm.

Um but uh again like Python is where our

focus will be. It allows us to be really

organized, loosely organized. Nice

flexibility there.

So, so far

based on these two slides, I'm showing

you that Python is interpreted, which is

easier and faster to work with. Um, not

faster to run, but faster to get up and

running because you don't need to

compile things. That's nice from our

perspective.

And it's also has very good flexibility

when it comes to organizing our

instructions, organizing our code can be

very loose in scripts, could be very

structured in in objects.

Okay. Okay. So generally no matter how

uh no matter what language it is um when

you process those instructions generally

things are going to be organized

even if it's in a script or if it's

object-oriented

um you're generally going to have the

very beginning of the program um kind of

setting up the input then the middle of

it really processing that and doing

something with that. So that's usually

like the bulk of the logic is in the

processing phase and then generally

you're producing some output. So that

could be like a model prediction, that

could be um a a graph that you've built

from your code um whatever that output

is. But generally it flows this way.

This is this is makes sense, right? Of

course there's input, you're

manipulating that input in some way and

then you're producing some output. I

think that all makes sense. That's a

very logical way to flow.

Um

now that's not to say that within this

processing step there may not be

um iteration like of course there may

may be times where we need to as part of

the processing kind of iterate and do

multiple passes of processing. Um so the

processing could be a lot. We could be

doing a lot. We could be doing a little.

Just depends on what we're actually

doing. So, if we're reading in some data

as the input um and then we're just

doing some simple um slicing and dicing

of it, that's some easy processing and

maybe producing a graph or producing a

metric, something of that sort, that's

pretty easy to do. But if we're training

a neural network or training a model,

the processing step can take a while and

it may, you know, be very iterative in

nature. So it just depends on what we're

doing and those instructions.

But no matter what, most of our programs

will flow in this way kind of input

processing output. It makes sense. It's

very logical.

So what are some principles that we

should abide by when we're writing our

code? So this this would really be for

any language, but of course for Python

that we are interested in. Um so

something we're going to be interested

in doing is um basically avoiding

repetition where we can. So instead of

having copy paste everywhere, we will

generally favor organizing our code to

some degree. Meaning we will utilize

functions where it makes sense and

objects where it makes sense to organize

things. And also instead of um having

very repetitive code, we will favor

using uh loop structures that can

iterate over um things many times

instead of us us having to write all

those out one by one by one. So we're

going to learn about these tools that we

have at our disposal, but they will help

us organize our code, avoid repetition

all over the place. One of the things we

want to avoid is having the same code

repeated all over the place. If if we

find ourselves doing that, we should

really put that code into a function or

maybe into an object so that we can

reuse it. So, we're really going to

favor like reusability of things,

re recycle, reuse, you know. So, we're

going to learn how to do that, how to

build functions, how to build objects.

But that's something we're going to

favor uh when we're when we're

programming. It's something you should

be on the lookout for. If you find

yourself writing the same code over and

over just in different spots, um that's

probably a clue you should organize that

into a function so you can just call

that function wherever you need to

rather than copying all that code. Okay,

so we're going to avoid repetition.

Now, the the reason we're going to do

that is to uh you know keep everything

simple. We want to make sure things are

clean, simple, understandable.

Um, we don't want to h we don't want to

have overly complex things that are very

difficult to follow. So, one of the

things that is going to be really nice

about Python is it lends itself very

well to being simple because it's going

to be so easy to actually read and

understand um, you know, understand

what's going on. But one of the things

that falls in line with this is like um

for instance naming things

appropriately. So instead of just

calling everything in our code like X Y

and Z if somebody comes along and reads

oh I see your code has an X Y and Z that

may not make sense. You know we would

want to be more thoughtful with the

names of our variables and names of our

function. So instead of XYZ, maybe we

would use something like name or place

or you know something appropriate to

identify this is what this is. So think

about that when you're writing your code

is try to make it understandable. Name

things that somebody else reading it

would understand what it is if they see

that name. So that's that's a mistake I

see a lot of people make when they first

start. It's okay like when you're first

getting started and practicing to name

things like X, Y, and Z. I think that's

fine. Or like ABC.

Um,

but does that make sense? Like if

somebody else was reading it, they see

XYZ in the program, that may not make

sense, you know? So, but if it has a

good name to it, you could say, oh, like

I see this is somebody's name that this

variable is referring to or this is um a

particular object that this is referring

to. Um, it's not just kind of an

abstract X or Y or Z.

Yeah, no spaghetti. Yeah, that's that's

what uh that's what a lot of people

refer to that as. Uh just sloppy,

unorganized, um hard to understand code.

One of the things that's great about

Python is it's naturally very

understandable. So like I don't think we

will have that issue as much as if we

had other languages, but it's still

possible.

So these are things we'll learn as we go

along. I'm just trying to get it into

your mind a little early here. Name

things appropriately is main one of the

main pieces of advice I can give here.

Um

so the next tip is to organize things.

This goes along with avoiding

repetition. So organize

um let's put things into functions.

Let's put things into objects where it

makes sense. If we know we're going to

reuse that um let's put it into a

function. And so we're going to learn

about how to do that. But generally this

is good practice if you find yourself

writing um uh code to do something and

it turns out to be um

it turns out to be uh something you know

you're going to reuse or it turns out to

be more than a handful of lines of code.

Generally you want to organize that into

a function so that uh it's clear

this is what this code is doing. This is

what it's responsible for. it's obvious

um you know that it's organized into

into uh that unit of work essentially.

So we are going to practice this. This

is something we're going to get good at

I think as we go along because we're

going to favor organization where it

makes sense.

Okay.

So readability. One of the things is

using meaningful names. I kind of

already mentioned that. The other thing

is using good comments. So, we're going

to learn probably today how to make

comments in our Python code, which is

going to be helpful to orient yourself

or another reader of it to, hey, this is

what this function does. This is what

this line of code is doing. Um, I can't

tell you how many times, you know,

people write code and then it they

themselves come back to it a week later

and have no idea what it's doing. That

happens all the time. It's even happened

to me. So, uh, comments are your friend

in that regard. and that um they don't

really cost you anything to put comments

in there um to say to to kind of

highlight this is what this piece of

code is doing and you can make a note to

yourself right within the code. That's

what comments are. They're basically

notes to yourself. Um

so we're going to learn about that

today. How to write comments and and

what that looks like in the code. The

other thing is indentation. you know,

Python supports uh indent like you have

to indent. So, that's not really going

to be an issue. Some languages don't

really support that, especially the

compiled ones. They don't enforce

strictly indentation. They enforce other

things like braces and and semicolons

and such, but um our our Python code

will be properly indented uh by

necessity because otherwise it won't

work. So, um, that's something we're

going to learn about too today is how we

indent things and why that matters.

We'll talk about that.

Um,

I see a question from Sherry. Is Python

a program that can be programmed with

simple language? Yes, it's very easy to

uh it it's Python is a very natural

language to program in because um yeah

it's very simple uh simple languages

used all over the place. I think it's

going to be really easy to learn. I

think it'll be really easy to pick up.

At least that's my hope and I think it

from my experience it is. As I said I

was someone who did that and I've worked

with many learners who've done the same.

So yes, I think it'll be pretty easy to

pick up, very simple.

Um, and then the other thing is we can

do uh we can find our errors very

quickly. Now, because this is an

interpreted language, we can run things

one line at a time and we we will

quickly hit errors

uh early on in our code if if we have

them. So this will be nice and Python

provides really good um error messages

um to say hey like this is what's wrong

with your code you should fix it this

way um essentially like giving you a

clue into what needs to be fixed. Um so

so this is something uh that we will

practice with as we go along is kind of

um finding errors and what to do with

them. Um, but because it's interpreted,

we will run across those very quickly.

Unlike with compiled language, which is

harder to debug because you basically

have to compile everything, hope that it

compiles. If it does, then you have to

run things. Um, it just takes longer to

get through that debugging phase. But

with the with Python, it's very quick.

You get a very quick feedback loop on if

your code's working or not, which is

nice. A lot of votes for C. I agree. AC

is the correct answer here. So the

interpreter is the thing that will

execute the code line by line. So it

doesn't do everything at once. It

actually goes line by line, which is why

you can stumble onto your errors quickly

because if you're going line by line

um and you have an error on this first

line, you're never going to reach these

other lines, right? You're it's just

going to show you this is where your

error is. it's on line 101 or whatever

it is and you know it's going to show

you where the error is. So it's going to

go one at a time and execute those. Um

it's not going to convert the code into

machine language. That's what a compiled

language would do, not an interpreted

one. Um and uh they do require an

interpreter. So D is just completely

wrong. It's the opposite of that. It

does require it. So the interpreter is

the thing that is executing the uh code

line by line. So what is Python in

particular? So it is a as we've already

seen an interpreted language meaning

that it requires an interpreter to

execute it. It's going to be executed

line by line by that interpreter. Um it

has capability to be object-oriented. It

also has capability to be scripted.

um which is just in relation to how it's

organized. One of the really nice things

is it is what we call dynamically typed

or what you would say dynamic semantics.

We will see what this means but

basically it means that we don't have to

declare what every piece of uh what

every variable or every piece of data is

inside of Python. We can let the

interpreter interpret that which is

nice. It makes things really easy to

work with. We don't need to say okay

this is an integer. This is a

floatingoint number. This is an array.

This is you know with a lot of program

especially compiled languages

programming languages you have to do

that because you have to tell the

compiler this is what this piece of data

is. But with an interpreter the

interpreter can as the name suggests

interpret that. It doesn't need to know

what everything is in terms of its data

type, which is which makes it really

easy to code. On the cons of that, it

can make it more prone to error because

you're not really enforcing types. So,

there is somewhat of a trade-off there.

But um for our purposes the dynamic

semantics make make it so that um the

interpreter can dynamically understand

what data is um based on how it's being

used which is great um for us like it

makes it just quicker to get up and

running and started and and working with

data. We don't need to declare what its

type is which is um static semantics.

Um now Python itself amazing programming

language that's used across many

different applications um such as data

science, automation, machine learning,

AI. It's also used in to build software

even um not sure if you guys know this

but there's um some really famous

software that's written in Python. Um,

one of the most famous is Instagram at

Meta is completely coded in Python,

which is it's over like 20,000 lines of

Python code, which is pretty amazing.

But um so of course it's been really um

heavily used in AI and machine learning

and such but it's also as a programming

language been used for other things like

more pure software applications which is

what makes Python really nice is it's so

simple so easy to learn. Um so for that

reason uh it is going to be great for us

to get started with especially if you're

coming in with basically no programming

experience. The other thing about Python

is it has uh as I said earlier like a

really big ecosystem uh meaning that

there's many different packages and

modules within those package packages

that do things already. So we don't

what's great about Python is we won't

need to reinvent the wheel on so many

different things like if we need to

build a plot if we need to train a model

and and use a specific type of model

that likely already exists in a package

somewhere. And what's great is they're

almost always open source meaning we

don't have to pay for anything. You can

just use it out of the box which is

fantastic. So there's within Python

there's so many ways to do things

especially in the AI and machine

learning world that we'll just borrow

those and use them in our own code um

which helps uh you know with um getting

up and running very quickly. We don't

need to reinvent things. We can just use

things that already exist um which is

fantastic. So that ecosystem really

benefits machine learning AI. Um because

they they already exist. We don't need

to spend our time rewriting all those

things. Um and so that's something we're

going to learn as we go along is like

how to install those, how to import

those, how to use those in our own code,

those those packages that already do

something for us. So we don't need to

come up with it on our own. we just need

to use it properly. Okay, so there's a

little bit of history. Python was first

invented in the late 1980s by a guy

named Guido Van Rossom in Amsterdam. Um,

where it gets its name is after the old

comedy series, you guys might be

familiar with it, the Monty Python

Flying Circus Show. Um, and so that's

where it's got its name. um you know it

was first created then but has since

taken on a really big role in the

especially you know I keep saying in the

AI community so much so that it has its

own software foundation that kind of is

responsible for maintaining it they meet

regularly they come up with improvements

um they come up with new versions of

Python

uh for example Python 3.14 just released

in October which is a major release. Uh

they hadn't had one in a while and that

one is uh 3.14. So it's kind of known as

Python.

Um which was a big milestone. Um but you

know they have uh they've had many

different versions over the years. It's

been maintained and developed by this

software foundation. Um and people are

actively working on it at many large

companies. So for instance, Meta has a

big group that is working on um Python

improvements. Microsoft as well, um

Google, all of those guys have groups

kind of working to improve Python

because they all use it. And so what

they typically do is work on it, open

source it, and then the community gets

to use those tools, those packages,

those tools, those improvements. Um so

it's it's actively um utilized across

many big companies actively uh

maintained by them or contributed to by

them. So that's that's really great. Um

you know Python was originally derived

from other language um other languages

uh as kind of a trying to find like a

mixture of some of the best of all

worlds. But its main like driving force

in why Python came to existence from

these other languages is it just its

ease of use. People really wanted

something like super easy to get up and

running and something really natural.

Um and so we will as we start learning

the syntax of it I think you guys will

understand why it's so easy. But um

that's that's what led to the

inspiration is just people wanted

something easier to work with not as not

as uh strenuous to kind of get up and

running.

What open source license is it? Um,

that's a good question. I think it's the

MIT license, but I could be wrong on

that.

You could look it up. If you go to

python.org.

Yeah, if you go to python.org, I think

it might talk more about what the uh

license structure is there. I want to

say it's MIT open license, but

I've I'm really not 100% sure on that.

Okay. So, what are some of the benefits

of working with Python? And these are

things you will experience as we go

along, but just wanted to call them out.

Um, the flexibility of it. As I said, it

can be really organized into

object-oriented or it can be loosely

organized into scripts. So, that

flexibility alone is really awesome. um

which has allowed it to power many

different things like um APIs, web

pages, full-blown applications like

Instagram, um chat, GPTs, like actual uh

AI, LLMs.

Um you know, it has so much flexibility

there to power so many different

applications.

Um probably the biggest benefit,

especially to us, is its ease of use.

Um,

uh,

oh, thank you. Some Tim just posted it.

It's the the GNU,

uh, public license. Yes.

Oh, never mind. It's a Python software.

It has its own. Okay, perfect. Thanks

for sharing that. Thanks for sharing

that. Yeah, I wasn't completely sure

which which license it was, but it is

open source. Um, and people do make

their own kind of derivations of Python.

But as I was saying, one of the benefits

of Python is how easy it is to learn. I

keep emphasizing that because it's true.

Once we get into it, you will see this.

I promise it'll be easy to learn, easy

to pick up. Um, and it's designed in

that way. Designed to be very minimalist

as a language, which is great.

um it has a lot of things that come with

it and it's kind of built into Python, a

lot of capability. So we call that the

standard library. It's just the things

built into Python. It has a lot of

capability out of the box. Um you know,

not only that, but it has a large

community that's developed so many

different packages that do things for

us, especially in the AI world. So

that's another great thing, kind of a

robust community developing these

packages that help us get things done.

Um, readability. So because the code is

so simple, it's also easy to read. So

you can usually read other Python code

and quickly understand what it's doing

which you know makes for easy um easy

understanding of other people's code

easy understanding of code in the

community and kind of almost like it's

selfdocumenting because it's so easy to

read. So that that simplicity that ease

of use lends itself well to being really

readable. You can usually just take a

look at the code, easily read it,

understand what it's doing, which is

great, like great for you guys learning,

great for taking a look at the demos and

examples that we will do. They're very

readable.

Okay. So why has Python really dominated

AI? So this is a valid question like

even so it's used for many different

things. It's a programming language. So

it can build application and I've given

you the example of Instagram and there's

many others um that are built off of

Python code. Why is it so useful for AI

in particular?

mainly

uh some of the reasons we've already

talked about mainly how easy it is to

use lends itself well for AI because um

that has allowed people to kind of

quickly get up and running and test out

their algorithms, test out their models

just really quickly with Python. That's

great. The other things listed on here

are certainly big reasons as well. So

for example, it has so many community

libraries, those those packages that um

have AI models and AI tools that we can

reuse that people have built these up

over years and years and years. Um so

it's to our benefit to reuse those and

not have to reinvent everything and we

can get quickly up and running with

those which would be great.

The other thing is Python, it lends

itself very very well to working with

data in general. Very easy to work with

data, very easy to load it in from

external sources, query it, work with

it, visualize it. Python is so adept at

that. Um, so that's what makes it really

nice at doing machine learning and AI

because so much of it is manipulating

data. So, um, for that reason alone,

Python is so popular in the AI community

just because of its ability to work with

data. It's so easy. This is something

we're going to really focus in on like

in our next course when we talk about

data science.

But, um,

just the ability and the power of it to

work with data makes lends itself well

to AI uh, capabilities.

Um, the other thing is I mentioned the

rapid prototyping. You can quickly build

a model in Python because the code is so

easy. So, and there's so many libraries

already can quickly prototype. Um,

it has obviously a big community around

it that's building out these packages,

writing documentation, maintaining it

from an open source level. So, that's

another reason it's very popular. Um,

Python's also used with other

technologies. So, it does have

capability to integrate with other

languages. So for instance, Python can

one of the most popular integrations is

Python can work with C and C++. So

sometimes that's necessary to integrate

with those to do certain things. Um so

Python has been extended to work with

other languages. So sometimes there's

other uh necessary support from other

like things in other languages that are

necessary to power something in AI. um

for example working with GPUs

and doing things in deep learning. Um

there's been a lot of integration with

uh working with um C tools. Now will we

do that? No, it's already been done for

us and some of these packages. But um

the pure ability of Python to do that is

really powerful and it gets taken for

granted honestly because you don't see

that it's underneath the hood and it's

abstracted away from you when you work

with those Python packages. But there

was a lot of work that went into it to

integrate it with other kind of other

programming languages.

Okay.

So as an example like I mentioned the

Instagram one. So Netflix for instance,

all of their recommendation is powered

by Python. So when you open up Netflix

or really any streaming service for that

matter, they're going to use Python to

deliver those recommendations and

produce those personalized

recommendations. Um Spotify as well for

like music. Um nearly all recommendation

algorithms are written in Python.

And in this program, we are actually

going to learn about recommendation

systems. So that'll be pretty fun way

down the road when we get into machine

learning. We'll talk about how do we

build a recommendation engine,

but um they're all done through Python

for for example. So really cool uh use

cases there.

So one of the things I wanted to address

is how AI itself is changing coding. So

you guys may be aware of this, but

obviously there's been a huge um kind of

explosion in generative AI tools that

can help write documents and write

emails and write text and all these

things. One of the things they can do is

write code. So um one of the big areas

where AI is changing coding is it's an

its ability to generate code for us. And

so um throughout this program like we

won't shy away from that necessarily

and I encourage you guys to use AI tools

as you see fit to help your own

understanding and help your own

productivity. Um

you know we still will go through the

fundamentals so you can understand it

but the AI tools can definitely be a

supplement to help. Um it's just that I

think you guys will understand it better

going through the examples that we do we

do together and so that when AI

generates code you will be able to

understand it and also be able to debug

it right because it's not always going

to be perfect. So that's always the

catch with AI is that you know it

doesn't always produce perfect answers.

Um but the at the very least we will be

able to you know debug things and

understand things better so that uh we

can catch those errors.

Um so obviously like AI is also besides

flat out generating it it's also

suggesting what should be there. So, uh,

some of the code editors really do a

good job at that, suggesting things, um,

picking up on what you should produce

next. That's going to be, um, very

interesting as we get into, uh, some of

the platforms that you guys will work

with to write your Python code. They

will have that ability. Um, so, uh, the

other thing is like there's some cloud

tools that, um, don't require writing

much code at all and they can just do

things. So, in other words, you can

power them by prompts. You're not really

writing code. You're just writing

natural language and then they do

something. Um, they generate the code in

the background and they execute

something. Um we will learn about those

things uh later on in the program

especially because we we will cover

generative AI in the future

um towards the end of our program. So if

you're wondering like are we going to

cover LLMs? Are we going to cover how

these things get generated? Yes. It just

will be um later on in the program.

Okay. A lot of votes for B.

Yeah, pretty unanimous on B. I think I

agree with it. Yeah, B is definitely the

right answer. So, all of the

recommendation systems which we will

learn how to build ourselves later on

are written in Python and um they uh are

machine learning models that make the

recommendations and that machine

learning is driven by data um and all of

that data is manipulated in Python

um and used to train uh models that do

the recommendations. That's all

happening in Python.

So, we're going to talk about getting

you guys set up on your own machine and

talking about the different development

environments we can use to actually work

with Python code. Um, before we go into

that, any questions about anything we

covered so far?

Everything's good so far. Yep. And you

know again if you have experience in

Python I recognize that it is going to

be a little slow in beginning. Um it's

mostly to get us really oriented to some

background around Python and get us set

up and then we will be doing you know uh

getting into the syntax and all that uh

coming up shortly. So we will actually

be learning Python specifics coming up

soon. But you know we're going to um get

everything set up first.

All right. So, let's continue then.

Thank you guys for that.

So, um it turns out that there are many

tools in the community for developing

Python code. And so, um you might hear

this word ID. It is short for integrated

development environment. This is a piece

of software that helps you write and

test Python code. So, and there's many

out there. There's a bunch on this list.

We are going to focus on a few options.

There's even more than what's on this

list, but we're going to focus on a few

options. These IDs are designed to

really help you write Python. They

provide many tools in the background

that make your life easier when you're

working with Python. So, for example,

they can provide syntax highlighting.

They can tell you when you have a syntax

error. Um, almost like a spell check for

Python.

Um, they can help you run Python code

right within the window. Um, they can

help you organize your projects. Uh,

they can do a lot of different things.

Um, and so there's many tools out there

that can do it, and it's really a

personal preference which one you use,

but in this program, we're really going

to focus on a few of them to to showcase

those options because they're very

popular options. Um, and then, uh, allow

you guys the flexibility to choose which

option makes the most sense for you. So,

generally, that's going to be mostly a a

preference.

um mostly a preference as to which one

you're the most comfortable with, but I

want to give you guys the option to uh

explore

the various options that are available.

Uh Roberto, is there one that stands out

as an industry standard? Um there's a

couple that you see like honestly the

two of them that we will study uh in

this coming up in the next few slides

are the industry standard which are

going to be VS code Microsoft VS code

and then Jupyter notebooks. So these two

are going to be uh ones that we will

study in particular and use throughout.

Um

so so yes we will those will be industry

standards. PyCharm's also very popular.

Um so I don't want to rule out PyCharm.

I know a lot of people who use it. So um

I would encourage you to explore PyCharm

as well if you want to but we are not

going to do that uh in in these slides

but um I would check it out and see if

you like it. Um it's another very I'm

putting a an asterisk next to it because

I think it's one of the more popular

uh yes uh yeah we're going to do

descriptions.

um requirements uh I'll try my best to

give those but honestly the requirements

will be given when you install them. Um

so the other thing I want to say is we

will have a couple options that don't

require you to install anything. So I'm

going to showcase those as well. So

there's a couple options that are um we

won't have to install anything because

they're going to be cloud-based.

Okay, I'll show you those.

Okay. So, but yeah, VS Code, I think VS

Code and Jupyter notebooks are are

probably the industry standard most

popular uh idees.

Okay. So, what we would recommend in

this program and the ones that we will

use the most uh throughout are going to

be these three. Visual Studio Code, also

known as VS Code, Jupyter Notebooks, and

Google Coll Collab, which is Google's

hosted

um Google's hosted version of notebooks

essentially. Um so

I will showcase each one of these and

give you some examples of how to set it

up and examples of how to work with it.

Um, and that's what we'll do over the

course of the next few slides and the

next uh bit of time is I'm going to go

through each one of these and kind of

show you what you would need to do to

get it set up. Um, now that being said,

excuse me, these two are ones that you

will install.

These two you would install locally on

your on your own machine.

And this one is um uh cloud hosted

by Google and it's free. Um all of these

are free but uh the first two VS code

and Jupyter notebook you would install

on your own machine. Collab you would

just access through your web browser. It

is hosted by Google. So that's an

advantage. You don't really need to

install anything. And for that reason um

sometimes we will favor Collab. Uh and

for other reasons too. Collab has some

really nice features if you've never

used it. Um, but notebooks, um, Jupyter

Notebook and Collab are very similar.

They're very similar. Collab just has

its own spin-off on on the notebook, um,

type of file that Jupyter Notebooks work

with. And it's um, like I said, kind of

cloud hosted. So, I'm going to I'm going

to walk us through each one of these and

explain to you what they do, what they

look like, and then we will um I'll set

up each one of them uh kind of in a live

demo so you guys can see. Um but uh we

throughout the program, it will really

be up to you which one of these you want

to use. There's no hard requirement to

use any one of them. It's really going

to be your preference which one of these

tools you want to use to work with

Python. Whatever one you feel

comfortable working with, that's the one

you should use.

All three of these are very popular in

the industry. So, you're not missing out

by using one versus the other. Um,

they're all very popular. Even Collab, I

know it wasn't on the screen, but it is

widely used in in the community and the

industry.

Uh, no system recommendations for

training LLMs. Um, no, because we don't

we won't really focus on that until the

end. When we get to when we get into

generative AI, we'll talk about that.

When we get into generative AI, we'll

talk about that.

So, yeah, we're not we're not focusing

on LM in the beginning. That's that's an

advanced topic for us.

What is my personal preference? Um, I

like Visual Studio Code. Um, personally

I that's what I use for my day-to-day

work is uh Visual Studio Code. I like

Visual Studio Code and I like Collab a

lot. Um, so you know, we'll talk about

this, but one of the advantages to

Collab is that it has free access to

GPUs, which is huge for doing things

like uh neural nets. Um, so we will lean

on collab quite a bit later on

uh later on when we um actually get to

deep learning and neural nets. We'll

because collab has free access to GPUs.

I'll show us that. It's it's really

nice.

And when you do anything with neural

nets, it usually benefits you to have a

GPU access. Um

so that'll be nice. But I usually do

most Python coding inside of VS Code. It

supports Python pretty pretty well.

What is more commonly used in the

industry? Um,

the two most popular are Visual Studio

Code and and Notebooks. Jupiter

notebooks.

They're both like you can't go wrong

with either one.

Those

two are really popular. Jupyter

notebooks and Visual Studio Code are

really popular. There's there's both of

those you would be okay with. Either

one.

Let me start with Visual Stu Studio

Code. So, um now Visual Studio Code

is a more general code editor. So, it's

actually you can edit lots of different

languages inside of VS Code. Um, so you

could do Java, you could do C, you could

do Scala, you can do Go, you can do all

kinds of languages are supported inside

of Visual Studio Code. So it's a really

fantastic product for programming in

general, not just Python. Um, it has

built-in terminal support. It has

co-pilot integrated into it, which is

nice for AI, like generative AI

assistance working with your code, which

is nice. Um, of course it supports

Python, which is what we are interested

in. Um, it has it has Python tools. I

will show us which ones we should

install as part of VS Code so that we

can work with Python files and

notebooks. Um, so it's it's a really

great code editor in general, which is

why I like using it. Um, but in

particular, it's pretty good at working

with Python. it it supports Python

pretty uh deeply. Um so and for that

reason VS code is really really popular.

But just keep in mind you can actually

use it for many different types of code

that uh that people write uh JavaScript

um Java as I said like many languages

are supported inside of Visual Studio

Code. So it's a more general code

editor. It happens to be really great at

working with Python though.

All right. So, I'm going to show us a

demo on setting up VS Code. Now, we are

going to do this for each one of these.

For Jupiter and for Collab, I'm going to

I'm going to do similar demos. So, um

don't worry, we'll get to those, but I

want to start with VS Code to show you

kind of how to get that set up and what

it looks like. Um, so where you can find

this demo

is inside of the demos that I mentioned

earlier in the reference material. So

I'm going to I'm going to jump over to

that. Let me show you guys.

So I'm back in the LMS. You guys will

want to download the demos. I think

somebody linked it earlier in case this

didn't show up for you, but we're going

to be inside of the demos and we're

going to do demo one for lesson one. We

do lesson one, demo one, which is going

to be the VS Code demo.

So, the main steps that we're going to

do is just going to be to point you to

where to install Visual Studio Code. So,

it is a it is an application is a free

application you can install on your

machine. Um, so,

uh, you will want to follow this link

that is within the demo file, this

code.vvisualstudio.com/d

download and download it for your

particular platform. So, if you're on

Windows, obviously, choose the Windows.

If you're on a Mac, um, choose Mac and

make sure that you choose the right, one

of the precautions is to choose the

right Mac platform. So, if you have like

an M1, 2, M3, M4 Mac, choose the Apple

Silicon

um button. If you're on an older Mac, um

then you'll want to use the Intel chip

one. Um

uh if you're on if you happen to be on

Linux, which I don't probably most of

you are not, but if you are, um you want

to download the right uh distribution uh

version.

But, uh follow this link first. So

that's the first step. Very easy step.

Just go to that site, pick your right

platform and uh go ahead and download

the installer. And mostly we will be

walking through the steps in the

installer. And then um I will show us

what it looks like once it's installed

and then show you a couple additional

steps that are actually not mentioned in

this file that I think are worth doing

to get you set up.

Uh yes, we will be doing Jupiter next.

Yes, we we'll we're going to be covering

VS Code, Jupiter, and Collab. I'm going

to show us examples of all of those.

Okay, let me ask you guys. Were you guys

able to get to the download page and

start that download and installation of

VS Code?

able to do that.

Any issues with that?

Okay. Yeah, it's just like any

yet I love I love the optimism

yet.

Uh already having both of them

installed. Okay. Yeah. No, if you

already have it installed, I mean,

great. I'll show so if you if you

already have VS Code installed, great.

You can sit tight. I will show you um a

couple of extensions that you'll want to

add for Python support

if you have it installed already. I'll

show us how you can use it with Python

in particular.

Okay.

If you already have it installed,

perfect. Looks like you have it

launched.

Still working on it. Okay. So, these

these instructions um uh show an example

of someone that would be on a Microsoft

platform um walking through the

installation.

Uh if you're on a Windows, you probably

want to create a desktop icon. You

definitely want to add it to your path.

And this just shows what's being

installed. So this is all the install

wizard on Windows. Nothing that exciting

there. So this if you follow all these

steps, you will have it installed. I

hope you have enough disc space. Uh I

don't think it's too big.

I don't think it's too too massive. I

forget how much space it takes. I don't

think it's that much.

I don't think it's that much. But um

yeah, hopefully you have enough.

So if if you don't

uh if you do not have enough disc space

um don't worry because we're going to do

collab which doesn't require you

installing anything. So you can always

use that option. All right. So if if for

some re let me just say that too just

even if if it's not a dispace issue if

you have an in any installation issues

no worries because we will work with

collab and Google that is going to be

cloud hosted that you don't need to

install anything you just need a Google

account

okay a free Google account

um so no worries at all if you cannot

get any of these things installed the

which are going to be Jupiter and

uh Jupiter and VS Code.

Where do we go? I haven't said yet. It

just I'm just making sure it's installed

for folks.

I'm going to I'm going to go over to it

in a second, but did we generally get it

installed and do you have it open? So,

if you once you get it installed,

uh once you get it installed, then open

it.

Yeah, you need to get it installed. Uh,

it should be this first. It should be

this link here.

Follow this link to get it installed.

Oops, I pasted the wrong link.

Let me find I'll copy and paste the

link. But yeah, take take a moment to

get it open. Once you have it open, just

sit tight

if you want to.

What does it say?

Yeah, feel. So, for you guys seeing the

co-pilot features, um,

click click use AI features. I think

that's okay. Yes. Um, you'll you'll

likely want copilot. Yes.

Click click okay on that.

That's the link, by the way, for the

download

in case uh we needed to get to it.

Okay.

So, I'm going to go over to VS Code

and show you what it looks like on uh my

end.

Okay. So, you should have something that

looks roughly like this. I don't have

anything open. I don't have any files

open. Uh just kind of have a blank

screen here. Um, but if you I would

recommend uh using the AI features if

you can. Um, I think that'll come in

handy later on.

Um, are we

comfortable uh moving forward? I want to

show us the extensions that support

Python. So, right now when you first

when you first install this, it does not

work with Python out of the box. We have

to install a couple extensions inside of

here to get it to work with Python. I'm

going to show us how to do that.

Don't worry about tuning any settings.

No, don't worry about doing any of that

at this stage. Don't really need to tune

anything. We just need to get Python

support.

Okay.

So, you guys with me on this main page?

You can use your corporate. Sure. Sure.

Yeah, you can you if you have it. If you

have co-pilot and want to use your

corporate, you can use that. That's

fine.

But you guys are with me on the main

page because I'm about to show us uh I'm

about to show us the extensions we need

to install to work with Python.

Okay, really important because this

isn't this is not in the documentation.

Um, no, no need to reinstall. Um, you

can I'll show you how to add that

through the extensions. No need to

reinstall.

You can add it as an extension. Yeah.

Okay.

So, let me ask you guys on the left,

do you see

this

little box icon that if you hover over

it says extensions?

Do you see that? you. There may be other

things here too, but at least that one

with the extensions.

Okay, so we do see that one. Okay,

so what we want to do,

no, I wouldn't I wouldn't uninstall.

That's okay because we're actually gonna

install Anaconda to get Jupiter. I

wouldn't un I wouldn't I would cancel

that if you can because you're going to

want that for Jupiter as well. I

wouldn't uninstall Anaconda.

I wouldn't uninstall. But I mean, if

it's already going if it's already doing

it, that's okay. We'll just reinstall it

later. All right. So, back to the

extensions. So, let's click on the

extensions.

Okay. So, do we see something like this

that has a search bar for extensions?

Do we see the search bar for the

extensions?

Okay. What do you think? We're going to

search for

Python.

Python.

We're going to search for Python. Yeah.

So, you are going to want to install the

official Python extension from

Microsoft. It is this one that has the

blue check mark next to Python.

Uh, so there now there are other ones

here,

but we just want the one that says

Python

from Microsoft. Do we see that extension

when you type in Python? Do we see that

one?

So just so it should just say Python. It

should be Microsoft.

Uh it's really popular. It has a lot of

downloads. Over 192 million downloads as

an extension.

It's from Microsoft.

Okay. Click on that.

Click on that.

And then you should see an install

button. It I already have it installed.

So it says uninstalled. Right here there

should be an install button. Install the

Python extension.

So out of 192 million installs,

really popular extension.

Are you guys able to install it?

You want to install that? It should be

pretty quick.

It should be pretty quick. It's not that

big of an extension.

So, what this does is

just the Python Sherry. It's just a

Python one. If you go into the

extensions and then search for Python,

it is just the one. It's just this one

that says Python and it's from

Microsoft.

Python blue check mark Microsoft.

You want that one.

And then you want to click on that one

and then hit the install.

Um, Roberto, is that for a co-pilot?

Is that for a co-pilot? I

maybe try closing it and reopening it.

Try closing VS Code, reopening and

retrying the install.

Um, no, we're not opening any folders

right now. We're not opening it. We're

just installing the extension.

That's all. We're just installing the

extension.

We're not opening any project folders.

just installing the extension.

Were we were we able to install that?

I know there's a lot by Microsoft, but

there should just be one that that says

Python.

there. So see how the name like this

name is this name here is eyesore. This

name is Python debugger. This one is

pilance.

Just the one that says Python.

Just that one.

That's the one we want. Only that one

right now.

Okay. Perfect. Perfect.

Okay, great.

Okay, so one more extension for you

guys. So once you install that one, I

have one more for you that you want to

install.

Are we ready for that one? One more we

want to install.

Okay, we're ready for the next one. So,

the next one you want to install

is the Jupiter extension,

which is the Jupiter.

It's this one. It's the very first one

here on my screen. So, it's it says

Jupiter

and it's from Microsoft.

Okay, we want to install that one.

Jupiter and it's from Microsoft. want to

install that one.

So, this one has 98 million uh installs.

You want to install this one.

Did you guys find that one? So, you want

to type in Jupy

Ter and it should be the Jupiter

extension here

that is uh from Microsoft.

So you want to install that one.

Great.

Now what does this one do? This

extension will allow you to work with

Jupiter notebooks inside of VS Code if

you want to.

So you Jupiter notebook has its own

standalone program which we will look at

next.

But you can open you can have those

files, those Jupyter notebook files be

compatible with VS Code and open them

and edit them and run them inside of VS

Code if you want to. So this extension

gives you the flexibility to work with

notebooks inside of VS Code. So you

never have to leave VS Code if you want

to work with notebooks. Um,

so this is a good extension if you

really want to work with notebooks and

stay inside of VS Code.

Yes. Uh when you Yeah. When you install

install an extension, it might it might

install a couple other dependency

extensions. Yes. But that's okay. Those

are required. That's okay.

That's that's that's okay.

All right. How do we feel? Good. Uh did

we get those installed?

Did we do were we able to get those

installed?

Okay, here is how we will test that it

all worked. So, we're going to do

something really simple.

Here's how we will test that it worked.

Let me go out of here and back to our

files.

So, out of the extensions, I just went

to the top button where it's the little

file um icon and um I am going to

um

go up to the very very top where it um

so you guys see on your VS Code window

where it says file, edit, selection,

view. I'm just going to create um

I'm just going to create uh a new

new file.

So, do you guys see that where where you

say file edit selection view? Click on

file and then click on new file.

You should see what I see on this screen

right here.

If you see

if you see Python and Jupyter notebook

then you know those are installed

correctly.

Do you guys see these options text

Python and Jupyter notebook?

Great. So what that means is we we can

now create those kind of files in the

future. We can create notebooks. you can

create Python files and VS Code will be

able to work with those.

If you don't see Python, that means your

Python extension didn't install yet or

you didn't install it. So, you want to

go back to you want to go back to your

extensions and make sure you installed

Python.

So go go go to this button over here,

the extensions,

type in Python,

and then make sure you install this

Python extension.

Okay. So, you're going to install the

Python extension and you're going to

install the Jupiter extension,

which is this, and install both of

those.

Make sure those are installed. If

they're installed and you still didn't

see that when you went to file um new

file,

if you don't see those, then um try

exiting VS Code and relaunching it.

Okay? Try exiting VS Code and reopening

it and seeing if you can make a new

file.

Okay? But it should be under uh at the

top file and then new file

and then you should see those options

Python and Jupiter.

Once you have those extension installed,

you may need to close out of VS Code and

reopen it to see that.

Okay, perfect. after you relaunched.

Okay, perfect. Yeah, you may need to

relaunch so that it can show the it can

show the extensions.

Yeah,

perfect.

Okay, perfect. So, that's set up for you

guys. So, um Perfect. It's set up for

you guys. Uh we will work with it in the

future, but just wanted to make sure it

was installed and set up. Once we start

working with Python, um I will show you

guys how to how to work with it. Um but

glad that's set up for now.

Uh what issue are you having uh Romero?

Is it not showing? It's not showing

Python or Jupiter for you when you do

file new file.

It's not showing those.

You may need to exit VS Code and reopen

it.

You uh Sil, yeah, you can you can make

one. We're not going to do anything with

it right now.

It's not going to you're not going to do

anything with it right now, but um

it's make sure you're searching for it

with a Y. It's J U P Y T E R.

You have to search. You have to So when

you go when you click on the extension,

search for JUP

Y. It should be the first thing that

shows up with Jupy Ter.

It's this Jupiter one from Microsoft.

I kernel I'll So the let me show us let

me show us that later. The kernel you

have to um you have to have a Python

interpreter.

So you may need to install a Python

interpreter to to be able to run the

kernel. So, I need to show us that. Um,

but I I don't want to get into that

right now.

Save what to

Oh, wherever you want. Wherever you want

on your own machine. It's up to you. It

doesn't really matter. Just wherever you

want.

It doesn't matter. It's up to you.

All right. So, what I want to do is uh I

want to take a break. Um because now,

you know, I said after two hours, we'll

take a longer break. Um so, we will now

we'll take a 10-minute break. Now, um if

you're still having any issues, um we

can try to get you set up at the end of

class. Um but we are going to set up.

So, coming up after our break, we're

going to take a 10-minute break. Coming

up after that, we'll we'll go and

install Jupyter Notebook. And then after

that, we will look at Collab. So, you're

going to have multiple options to run

Python. Not So, if this wasn't working

for you, that's okay. We'll try a

different route.

Okay? I will try a different route. Um I

I know Collab will work for you because

that is hosted by Google and really easy

to get working with. So at the worst

case scenario, Collab will work for you.

I know it. Um but we'll try to get

Jupyter Notebooks installed for you as

well. But if you're having issues with

VS Code, let me know at the end of

class. We'll try to get you set up,

okay?

You're still having issues with it.

But um what we're going to do right now

is take take a 10-minute break.

So let's try to be back um in about uh

10 minutes. Let's call it an even um

let's call it an even

uh what will we be covering? Um

installing the other installing the

other um Python setups. So Jupyter

notebook and working with collab. And

then we will get into the basics of

Python's the syntax. So we're going to

talk about indentation, identifiers, um

maybe if we have time, basic variable

types, data types. Yep. So we'll get

into Python.

We will get into Python today.

All right. So let's jump over to

uh Jupiter notebooks. So um what's so

special about Jupiter?

Well, it turns out that uh Jupiter is a

platform for running what are called

notebook files. So obviously we just

installed the Jupiter extension in VS

Code which will allow us to run

notebooks in VS code but Jupiter has its

own notebook platform and that's what

you will install in this setup. Um,

notebooks are special. They are um

really great um Python code files that

give us the ability to execute isolated

what are called cells of code. So we can

run one cell at a time and test and

debug the execution of that single cell

without affecting any of the other

cells. So, um, notebooks are great for,

uh, running code live and interactive.

When we do a lot of our demos in this

program, they're all going to be in

notebooks. Um, so that we can kind of

run things one cell at a time. Um,

uh,

no. So without notebooks you either have

to run you run like a Python script like

a Python file um which is a py file and

usually you have to either run that

through a debugger or run the entire

script at once. You don't really get

code isolated into individual cells

which is really nice with notebooks. The

other thing is notebooks are easily

sharable.

So you can share a notebook with

somebody else and they can open it and

see all of your inputs and outputs in

the notebook which is really nice like

all of the outputs get saved into the

notebook. Um which is nice. So and

notebooks uh especially in the Jupiter

platform are going to have all the data

science libraries available to them. So,

uh, if you're people usually love doing

notebooks for working with data, um,

really easy to work with data inside of

notebooks and and build things like

plots. You can display your,

uh, you can display your graphs really

easily inside of the notebook and then

share your notebook so other people can

see your graphs. Um, so notebooks are

really awesome like interactive

environments for running code. Um and we

will favor notebooks uh as our primary

way of running code throughout the

program. Now where you open those

notebooks is up to you. You can open

them in VS Code. You can open them in

the Jupyter notebook platform. Uh

you can open them inside of Collab and

run notebooks in Collab. Uh notebooks

are very very popular.

Why isn't running in notebooks the

default? It's because uh not all code

runs inside of cells. Like applications

are not going to be well suited for

notebooks. Like Instagram is not running

in a notebook. Uh it's more structured

into actual Python files and actual uh

more structured programs are going to be

not in a notebook. Notebook is more for

prototyping and debugging and uh

executing small chunks of code to test

it out. It's not for writing larger

programs like an like a

an LLM application like a chatbot would

generally be in not in a notebook. It'd

be in like a Python file.

Uh cells versus class objects. So cells

are just small uh think of them as small

little environments to execute our code.

Um class objects are actual chunks of

code that define an object. They're

they're different things.

Yeah, different things. We'll we'll

learn about objects. Um and we will

certainly see what cells are as we go

through. I'm going to show you an

example of a cell coming up when we

install Jupiter.

But uh let's talk about let's uh go

through the installation of Jupyter

notebook so you can see what a notebook

looks like. I think that'll be helpful

to orient.

So let's go over to that demo. So this

is going to be demo two

uh demo two inside of um uh lesson one.

So we're going to go over to that.

Everybody has this one. Okay, perfect.

Okay, so you're going to follow this

instruction. Now, what this is going to

do is first

um

No, this has not this is not going to be

anything to do with VS Code. This is

going to be a different platform. This

is going to be Jupiter.

Where is this? This is the

This is the demos.

This is uh demo two inside of that demos

folder that we said uh

to to uh grab all the demos

from your LMS.

Does anybody have that uh demo 2 PDF

they can upload? I I think somebody

uploaded all of them earlier, but if you

have demo two, want to upload it real

quick? I don't have the PDFs.

if somebody wants to share that.

So there so they're different. Um VS So

what I was saying is you can open

notebooks inside of VS Code and the

thing that allows you to open notebooks

in VS Code is the extension.

So yes, if you're going to work with

notebooks in VS Code, you need the

extension installed. But you can use the

standalone Jupiter platform

to work with notebooks. It's up to you.

If you like using VS Code,

um if you like using VS Code, you can do

it that way. If you like uh the Jupiter

platform, you can do it that way. It's

up to you. It's just a preference. I'm

giving you guys options. That's my goal

is to give you options and let you guys

choose what you're most comfortable

with.

Okay. And we're and we're taking time to

do that now in the beginning of the

program, right? Because we're going to

be doing a lot of Python examples coming

up as we start learning Python. So, it's

it's valuable to spend that time now. I

know it can seem a little slow, but I

promise it'll be worth it so that you

guys have options for running your

running your code.

Yes. Thank you guys for uploading those.

appreciate it. Those are the demos you

want to uh follow along with.

Okay, so the first step here is going to

be to install Anaconda. Now, you may be

wondering, what is Anaconda? I thought

we were talking about Jupiter, and

that's a valid question. Anaconda is a

what's called a distribution of Python.

So Anaconda is a program a software a

collection of software programs that

give you a version of Python with a

bunch of packages

uh with a bunch of packages already

installed.

Um and then

uh one of those is the Jupiter package

so that you can run Jupyter notebooks.

And what Jupyter notebooks will be

is a uh basically a web browser

application that will open up a notebook

editor in your web browser. So that's

ultimately what we're going to do, but

we are going to install it via the

Anaconda distribution

uh via the Anaconda distribution of

Python.

So that's where we're going to start is

with the initial download of Anaconda.

Oh, it's no no skipped registration.

Okay, let me let me uh open the link.

I think there I think there's a way to

find it without having to do the

registration.

There's a way to get to it without

having to do that. I'm going to find it

real quick.

Oh, you can't. Okay. So, if you can't

install it, that's okay. We will be able

to work with notebooks in collab and you

can work with notebooks inside of VS

Code. That's fine, too.

Yeah, I'm getting I I'm going through

the registration process so I can um I

can show you that install.

Okay, let me share my screen.

Did you guys get to once you go through

the like setting up your account, do you

get to this page?

Do you get to this page for those of you

going through? Yeah, that looks right

for you, Ashish. That looks right.

Do you guys get to this page though when

you get through your like account setup?

Okay, you got to this page. Okay, so

then choose your correct Windows or Mac

down. You want to be over here on the

left. You want to do Anaconda

distribution.

This is what you want to do. So, choose

the right one. And if you're on an M1,

M2, M3, you're going to do the silicon.

If you're on an older Mac, you're going

to do the 64. And then obviously, if

you're on a Windows, you should be

clicking over here to do Windows. But

you want to do the Anaconda

distribution, not Minion. Okay. So,

click on the installer for Anaconda

distribution.

Okay? And then let that install. Now

while that's installing let me explain

something about the difference between

uh I think it was asked earlier what's

the difference between um Anaconda

uh as the default Python. So Anaconda

as I was saying earlier is a version of

Python that has a bunch of data science

and machine learning packages already

installed for you. So uh it comes with a

bunch of packages that are already

installed. So if you use that Python

um that Python has a bunch of packages

built in with it that you don't need to

go out and install. So, Anaconda is a

very popular version of Python for

people to install that are working in

data science, AI, ML. Very popular

version because it already comes with a

bunch of packages that you would use for

manipulating data for doing machine

learning or doing anything in AI. So,

it's it's a very um popular

distribution. It also comes with

Jupiter, which is why we wanted to use

it because it comes with the notebook

capability out of the box.

Okay. So, I'm going to launch. So, when

this is done installing, you want to

launch the program that gets installed

called the uh Anaconda Navigator.

So, it should install a program on your

machine called the Anaconda Navigator.

Do you guys have that? Did anybody get

through and and you have that program?

The Anaconda Navigator.

You don't need any advanced ones. You

don't need any advanced options.

Just the just the defaults. All the

defaults

should be good.

Still downloading. Okay. I'm going to

show you

I'm going to show you what the navigator

looks like once you once you have it.

That's okay if it takes a little bit of

time to download. That's okay. Um,

basically once you download it, um, you

just have to click a couple more buttons

and then you can access Jupiter.

Okay, let me share my screen and show

you what you like. Once it installs,

this is what it should look like. It's

okay if it's taking a little bit of

time.

You should have something that kind of

looks like this, which is the um

dashboard that has the different

programs available to you to you.

Um

do you guys see something like this? If

you have the navigator,

do you see something like this?

which is the which is the like when you

open the navigator program, you should

see something like this that has a bunch

of different um

you do. Okay.

It's if it's taking a little bit of time

that's okay.

Yes. Na Anaconda Navigator is how you

launch Yes.

Anaconda Navigator is how you launch it.

So yeah, you want to open that. Now the

the whole reason to come here

is so we can launch Jupiter notebooks.

So we can launch Jupyter notebooks. Um

this is the program we ultimately want

to launch. This is going to

uh allow us to open notebook files, edit

them, run code cells. I'll show you what

a notebook looks like in a second. But

but once you have Anaconda installed,

open the navigator and then launch

Jupiter notebook. It's just one extra

step. Launch the Jupiter notebook.

What that should do is launch the the

notebook.

Uh it should launch the web browser of

your like whatever you have as your

default web browser. It should open the

notebook program in every in your web

browser. So if it's Chrome, Firefox,

Edge, whatever your default web browser,

it's going to launch the notebook

program in the browser.

Okay.

So, I'm going to launch it and then I'll

show you what it looks like. Again, if

it's taking you a little bit of time,

that's okay. Whenever it's done,

how did I get to these icons? Just

launch. Do you have the Anaconda

Navigator program?

It should have got It should be

installed.

Open the Anaconda Navigator program.

It should have been installed with the

Anaconda installation.

All right. Was anybody able to get to

this the Jupiter this? So, it should

launch in your browser. Anybody

able to get to that?

Fantastic. Fantastic. I'm glad some of

you guys are able to get to it. And if

it's it's not yet, that's okay.

Remember, when it's done installing,

you're going to go to Anaconda Navigator

and then

uh launch Jupiter Notebook. That's

That's what you're going to do.

That's okay, Roberto. It's okay.

All right. I do want to I want to show

you guys a notebook. I just want to show

you what it looks like. What I'm going

to do is I'm going to

um open a notebook by going to new and

then Python 3 notebook. So you can open

a folder, you can open a terminal, you

can open a text file. I'm going to open

a Python 3 which is a notebook. You so

it's a it's a a notebook powered by

Python.

So I'm going to click on that which will

launch a new notebook in a new tab.

And here I am in the notebook editor. So

now I am in a notebook editor screen. So

if you go when you first launch Jupiter

you you can navigate to notice that that

notebook got created here where I

currently am on my machine. I could

navigate to I could navigate to

documents and then I could um you know

create a new file there or I could make

a new folder here and and do it that

way. Um but uh I am uh just editing this

notebook right here within this um

current folder that I'm in.

Okay. So, do you guys remember when I

said that code gets executed in a in an

isolated cell?

Do you remember that?

Um,

this is what a cell looks like. And you

can make new cells by hitting this plus

button.

So, you hit this plus button over here,

you can make new cells. So, if you hit

plus++,

I'm making a bunch of cells.

Now, what's really cool about cells?

Yeah, Tim just discovered this. What's

really cool about cells is you can

change them to be text or code. So, if

you change it to markdown,

I can write markdown text in here to say

this is my notebook. And then if I run

this, it's going to display as text.

So if I run that cell which uh when I'm

editing it I can click run and it will

render that as text because I changed

the cell type to markdown. Markdown is

just a flavor of text style.

So otherwise we can write some Python

code. Now, what I want you guys to type,

I'll type this in the chat to verify

everything is working is I want you to

type print

hello world.

I want you to type that

inside of a cell

and then

and then hit run.

And it should run that code.

And you should see you should be able to

see uh you should be able to see that

Shift enter. Yep. You can whenever

you're on a cell, you can hit shift

enter. It'll run the cell.

You can That's okay. You can always You

can go back and watch the video. So,

this is being recorded. You can go back

and watch the video. I I know it's a

little frustrating. It's still

installing for you, but go back and

watch the video. And I definitely

encourage once it's installed to go back

and try this, which would be just

launching your Anaconda Navigator

and then launching Jupiter.

If you don't have, by the way, if you

don't have Python 3, um you may need to

uh exit your navigator and reopen it.

Okay, you may need to exit your your

navigator, reopen it so that you can

launch Jupiter again.

Were you guys able to run this in a

cell? For those of you that have Jupiter

running, were you able to run this?

Nice. And it worked for you. Okay,

perfect. Perfect.

So this is what I meant by this is an

isolated cell. So notice that we can run

this

and it doesn't affect

um

Sure. Sure. I hear you. I I hear you.

Update the doc. Uh I can How about I

post it in our um our Slack channel? By

the way, do you guys have access to the

to the Slack channel?

Okay, I can post it there. I can post

the instructions to get there.

Okay, I can post it in our our uh

cohort's uh channel.

I hear you. You don't want to search for

our video. I I hear you.

Uh I don't have the link on hand, but

you can get to it through the LMS.

So if you go to the LMS and go to

uh there should be

um

there should be a link to get to it

within there. It's should be like over

here on the right.

I don't have the link I don't have the

link to it off hand. Yeah,

but there should be a way to get to it

from the LMS.

Yeah, there should be a banner here. I

don't know why I don't have it, but

should be there.

Okay. So, if you're just getting things

installed, how do you get to here? Um,

you open the navigator.

Open the navigator.

Syntax is hello world.

It's just inside of it's just that print

hello world.

Um, open the navigator.

Open the navigator which looks like

this.

Let me share my screen.

Okay. Open the Anaconda Navigator that

got installed.

Then click launch on the Jupyter

notebook program. So you should have

this at least. You may have other ones.

Click on this launch. Uh click on this

launch and then you can launch the uh

Anaconda Navigator.

Okay.

If it if it's a little stuck, that's

okay. We're going to move on. We're

going to go to Collab, which can run

notebooks as well. So, if it seems a

little stuck, that's okay. I will post

in our Slack instructions on how to run

this.

That's okay.

All right. But what I wanted to do,

what I wanted to do before we move on to

collab is I just wanted to show you I

wanted to call out a couple things about

notebooks.

Um

is that uh a couple things about

notebooks. One is that notice that these

cells are very isolated. Whatever I put

here

um does not affect what I had before. So

I can add numbers like that and it can

um compute that and this does not affect

this. So so this is why notebooks are so

great is you can document the notebooks

with with mixing in text and code like

we do here. Um you can run code in its

own cells.

uh you can run code in its own cells and

then you can um have that very isolated.

So I could jump down here and run

something and that doesn't matter that

there's nothing here like it doesn't

need to be in order. I can um you know I

can uh run stuff out of I can overwrite

this

um and run that and it produces the

output. Uh

so you know many things we could do uh

inside of notebooks that are really

fantastic for just quickly prototyping

and running Python code inside of cells.

So it's very nice that way. So notebooks

are notebooks are really nice. You can

also like I could share this file. So

this produces a file on my machine. Uh

if I go back to the navigator

um

Anaconda navigator Jerry Anaconda

navigator

um

can you read value of variables from

another cell? Uh you you have to store

them into variables. So I could I could

call this uh x

and then I could refer to x later.

We'll learn about that. We'll learn

about that with variables.

But yes, you can kind of do that with

variables.

All right.

Um,

what do the numbers after?

Which numbers? these the ones in the

brackets.

Oh th so those are which cells we've uh

executed. So I executed this one first.

So it it's it's number one. And then I

executed um

uh this I think I did this second. So it

it's text. It doesn't really get one of

those. And then I did this one third.

And then I did um I think I did this one

fourth and then it got overwritten with

the fifth. So it just tells you like how

many executions you've done and what is

which number execution that was. Then I

did this one sixth.

It just keeps track of your executions.

Okay. So somebody asked about a kernel.

What is a kernel? So uh the kernel is um

the kernel is basically the interpreter.

So it's the thing that that the kernel

is just a a um a copy of the interpreter

that the notebook is attaching to in

order to run. So the notebook can't run

anything because remember a pi Python

needs

Python needs an interpreter to run its

code. So in notebooks we basically

create like a virtual copy of the

interpreter called a kernel. Um and you

can actually have many kernels based on

your um your base interpreter. So what's

nice is Anaconda

um Anaconda comes with

uh an interpreter for you and then you

create kernels that are virtual copies

of that um that are virtual copies of

that uh interpreter so that you can run

your Python code against it. Remember

you need an interpreter but notebooks

attach to kernels. Kernels are like

virtual interpreters.

Um, and you can have many kernels based

on the original interpreter. So the

kernel is literally just think of it

like the computer that's powering the

notebook. That's all. It's just the

compute engine that's allowing you to

execute your Python code. So every

notebook has an associated kernel.

And what's interesting is if you restart

your kernel, you lose all your data. So

all of these outputs that we have um you

would lose if you restarted your kernel.

So if I restart um now I like since I

restarted this is not going to know what

x is. So if I try to print x again it's

going to say I don't know what x is

because I restarted my kernel. I lost

all that data.

But I can redefine it. And then there it

is. And notice that my iterations

restart

um my iterations restart when I uh

restart my kernel. So if if I go back

and restart the kernel again

and now if I run this, this will be

first. So notice how that restarts to

first. This will be second. This will be

third.

Try restarting. I don't know what that

is.

I don't know why that

Yeah, choose the Anaconda. Either one.

Choose. You want to use Anaconda as your

But what that's saying is what do you

want to use as your interpreter to to

build your kernels. So, choose one of

those. That's fine.

So, yeah, Anaconda requirements, laptop

requirements. Um,

you need a little bit of you need a

little bit of RAM. Uh, you need a little

bit of RAM to run the notebooks because

you need some memory. Um, you don't need

a lot of it though. I'd be surprised if

you didn't meet the requirements. It's

not that much, but you do need some.

I'm not sure the exact. I'd have to look

that up on the Anaconda website.

Oh, it must have been full to start with

or pretty full. I'd be This doesn't take

up that much space, I don't think.

Was it pretty full to begin with?

I would assume. I don't think this takes

up that much space.

Uh, what I want to do then, I want to go

over to collab. Okay. How do we feel

about the notebooks? I maybe if it's

still installing for you, give it a

little time. Open the open the

navigator.

Let's try Coll. I guarantee you Collab

will work for you if you're still having

issues with if you're having issues with

Jupiter.

No worries. Let's just try collab. I

promise that will be a lot easier, be a

million times easier, I think, than than

working with Jupiter. Okay, great. So,

we will continue then.

All right, let me jump over to our

final

um

demo with setting up a collab notebook.

So I'm just going to jump into doing

that on in the interest of time.

Uh

so

would you be taking up additional

sessions too other than uh so we're I'm

going to be the instructor for all of

the courses in this program. So you're

you're stuck with me

for all of those. Does that make sense?

like all of the all of the uh AI

engineer program uh courses.

Yeah. Yeah. Stuck or be excited. It's

going to be one or the other. Probably

not an in between feeling.

Hopefully. Cool. Hopefully. Hopefully

good. Yeah. Like I said, I've taught

this many times. Uh, I think it would be

uh I think it'll be good.

You are stuck. Okay. Well, we're going

to get you unstuck with Collab. I would

not worry about getting Jupiter set up.

If it's not working for you, we're going

to ditch it and we're going to use

something else that works. I promise

it's not a I promise getting Jupiter set

up is not that important relative to

getting at least one of these options

that works.

So, if it's not working over on Jupiter,

I'm not worried in the slightest about

it because there's going to be plenty of

options to run run Python code. In fact,

we're going to do one next which is

going to be with um with Collab. So,

we'll do that. Um so, let me jump into

that. Let me share my screen here.

Um,

learning a lot already. Great. That's

great. Glad to hear that. Thank you.

Okay.

Thank you guys. Appreciate it.

All right. Let me go to the demo. Demo

three.

All right. So, what you want to do

essentially is to go to this website,

um, which I have here. I'm going to, uh,

paste it in the chat. Um, so what you

want to do is go to Google's website for

their Collab platform, uh, which is, so

Collab is a, um, notebook platform that

Google hosts. So, you don't need to

install anything. You just go there in

your favorite web browser, log in with

your Google account. In fact, I don't

even think you need to be necessarily

logged in. You can in order to save your

notebooks to your drive,

but um you go there and you basically

open up a notebook and start working

with it right away. And it's fantastic.

Their notebook environment already has a

lot of packages installed into it for AI

and machine learning. So that's f that's

really great. Um

uh once you get to the page um you

should log in though if you have a

Google account. Only reason I say that

is because it will save your notebooks

to your drive automatically so that you

it will automatically save your

notebooks just like as if you're working

in a Google doc. So that's great. So

that um it saves your work

automatically.

Um, so please, you know, I would

recommend getting a Google account if

you don't have one for free. Logging in

using Collab is completely free.

Um, so it's a fantastic platform. Um,

when you go to that site, uh, assuming

you've logged in, you want to click on

that lower left blue button where it

says new notebook. You can see it in

this screenshot. And I I'll open up one

in a moment on on my screen. But do you

guys see this screen right here that's

in this screenshot that has the new

notebook on the bottom in the lower

left?

No. From that site, what do you see?

Oh, so you're already in a notebook. It

you're already in a notebook. Like it

says, "Welcome to Collab.

Oh, okay. So, it already opened the

notebook for you. Okay, that's that's

fine. That's fine. I'll show you uh I'll

show you what that looks like. That's no

problem. That means you're already

inside of it.

Okay.

Okay. So, then we're pretty much in the

notebook environment and we can start

running code. Let me let me hop over to

Collab and show you guys what it looks

like.

Let me stop sharing that and jump over

to

collab here.

Okay.

So, if you're in the welcome to collab,

um that's fine or you can start a new

notebook. Let me assume that we've

opened up welcome to collab. So, you're

in this screen. What you want to do if

you're in this screen is just go to go

up to file and do new notebook.

Just go to file, new notebook in drive.

Just do that. File new notebook

and this will create a new notebook for

you which will start fresh a blank

notebook.

Okay.

Were you able to do that? If you guys

were folks able to get here to this uh

blank notebook

one way or the other, you clicked the

blue button to hit a new one or you went

up to file and did new notebook.

Yes. Okay.

So, there we are. Without doing all the

Jupiter install steps, we're in a

notebook. Look how easy that was, right?

So, why didn't we just start with this?

Um

so yeah so we're in Google's notebook

platform uh which is a fantastic

platform and uh what's great about this

is you can export these now these are

pyb which is which is uh if you're

curious what that extension means it's

short for interactive python notebook

okay IPIB

so these are the files that you can open

in Jupiter if you have uh if you notice

when you open up Jupyter notebook book

earlier it was a IP YMBB

um inside of VS code when you work with

notebooks they are IP YMBB so IPMBB is a

notebook file and it can be opened in

any one of these three platforms right

the Jupiter from Anaconda the uh VS code

can open IPMB and you can also upload

your own notebooks here if you have them

on your own machine you just go to file

upload notebook and And then it will

open up a box where you can choose which

file on your machine to upload. So you

can upload your own notebooks, which

will be uh great when we get into um

demos. We have demo notebooks for you

guys that we'll work through with our

code. You can upload those into Collab

and work with them directly inside of

here.

So let's try running something. Let's do

the print

hello world.

So, um, you want to type that in and I

can paste it in the chat for you guys

and then you want to you want to hit

either shift enter or this play button

right next to the cell.

Okay, so that might take a moment

because it's booting up your uh your

kernel.

Uh, but then it should run and you

should see the output. Now, this is

going to look very similar to Jupiter,

just slightly different.

We're inside of Collab

and we started a new notebook.

We're just inside of a blank notebook

for now. And we uh are just within this

first cell and I I'm doing hello world.

Were you guys able to run that?

We didn't. But we could we could open a

notebook in VS Code because we installed

the extension. We did that. Remember we

installed the Jupiter extension. So we

can open notebooks in VS Code and we can

run them there. I just didn't show that

to us. Uh we might do that later down

the road, but you do have that

flexibility to run things there if you

want to.

Okay. One thing I want to show you guys

that's really cool. So, um, one thing I

want to show you is if you go up to

runtime

and then go down to change runtime type.

Do you do you guys have that? Change

runtime type. So, if you click on

runtime

and then change runtime type.

Do we have that? Click on that. Click on

change runtime type.

And look at our options. We can choose a

GPU for free.

So we can swap over to a GPU kernel

which is fantastic for training deep

learning models and we can use that GPU

for free. This is one of the reasons

that uh Collab is so amazing is they

give free access to a GPU. So if you

don't have one on your own machine um

you can use the GPUs from Collab for

free.

Yeah, go ahead. I mean, there's no

nothing wrong with it. So, uh, what it's

going to ask you to do is, uh, terminate

your current kernel because you're

connected to a CPU basic kernel. It

wants you to disable that so you can

swap over the GPU. Click okay. That's

okay.

All right. And then we are now um, we

click save. And that will swap us over

and connect us to a GPU. So, if you how

you know that you're connected to a GPU

is if you go over to um

if you go over to this box on the right.

Do you guys see that one where it says

RAM and disk? If we click that,

it will show us our resource resource

usage. And you should see GPU RAM

available of 15 gigs.

So, you have 15 gigabytes of GPU RAM

available that you can use.

So remember, you just click this RAM

um

you should click this RAM uh

uh

sorry this RAM and disk.

Roberto, did you swap over the runtime

to

uh Yeah, it should pop up. Okay, then

you should be able to click on this

Yeah, it might take some time to connect

to one because what Google has to

allocate one to you um and then it has

to like connect it over the cloud. It

can take a minute. Yeah, it can take a

minute. It has to allocate one to you.

So the question is which one is better?

Um,

so for the vast majority of things,

the CPU, the standard CPU runtime, which

is the default, is going to be better

for the vast majority of things. The

only time the GPU is really going to be

beneficial is when we start doing deep

learning and training neural networks,

then the GPU will be really beneficial.

It will speed up the training time by a

significant amount.

I can tell you like I trained a uh

neural network for images for image

recognition. Uh it took me it took two

hours on the CPU and then when I swapped

it over to GPU it took less than a

minute

took less than a minute and it was

taking two hours on the CPU.

So yeah, training neural nets on a GPU

when we get to that is going to be

beneficial. So if you're not using

Collab right now, that's okay, but in

the future when we get into deep

learning, you're likely going to want to

use Collab to swap over to the GPU for

free.

Now, they do rate limit you,

so it's free, but you can max it out in

a day and then they cool you off for 24

hours, which I have I have done, uh,

unfortunately. So, like, if you max out

that RAM and you use it too much in a

24-hour period, they will, uh, not allow

you to connect to a free GPU for another

24 hours.

So, I doubt you'll run into that

situation, but I have before

if you're just if you're just using it

so much.

No. So, unfortunately,

uh, no.

So, unfortunately, no. You cannot, um,

Collab doesn't connect to your local

resources. It only it does the cloud

Google's cloud resources. So no, you

can't use your own through collab. But

yes, you could use your own GPU through

VS Code. I will show us how to do that

later when we get into deep learning.

I will show you that later.

We we won't need to worry about that

now, but later on, yes, that'll be

important.

Uh it doesn't show GPU. Make sure you

swap over the runtime to go to change

runtime type and make sure you pick GPU.

Make sure you go away from

No, you should use Collab. I wouldn't

You don't need to buy anything. You just

use Collab. Just use Collab for sure.

Collab's free. It does everything you're

going to need for the class.

Yeah, I I highly advocate for Collab. It

So, by the way, if you're curious, like

Collab came about because Google wanted

the the machine learning research

community to have access to GPUs for

free to um develop like machine learning

and uh deep learning models. So, uh it's

been around for a while. I remember

using Collab um probably almost 10 years

ago and it used to be it used to be very

lucky if you got a GPU. You used to like

you used to have to click and then hope

that you would get allocated one and

sometimes you wouldn't and I would sit

there and have to refresh and try to

hope that I would get a GPU but now it's

it's like readily available which is

fantastic.

No, you're But you're joining it at a

good time because I'm telling you, the

GPU was very difficult to get. I would

always try to switch over to that and I

would rarely be able to.

So,

pretty good. But like I said, like if

the CPU is perfectly fine for everything

we're going to do, except when we get

into deep learning, you're going to want

to swap that over to GPU. But that's

going to be for we have a while till we

get to deep learning.

We have a lot to learn between now and

then.

Okay.

What do we think? Do we like collab?

We're comfortable with it. Feel free to

use it. Feel free to use VS Code. Feel

free to use Jupiter. Whatever you want

to use, okay? There's options, right? I

hopefully you have options that work for

you. Um they are all used in the

industry. So you're not missing out on

if whatever you use, people use it of

these three people use all of them.

So feel free to use whatever is easiest.

Yeah, I can show that real quick. Yeah,

let me go back over to it.

I'm going to be real quick on it though

because I want to make sure we get over

to our other material.

Okay, let me show you how you can run a

notebook. Let me show you how to run a

notebook. So I'm going to go to file,

new file, and open a Jupyter notebook.

Okay. So it'll open a new. Now notice

notice the extension of it

is

MB. That should be no surprise. That is

the universal kind of interactive Python

notebook file.

Okay. So the biggest thing you have to

do when you open a notebook in VS Code

is you have to

tell it what kernel to connect to. So

have to go to select kernel

and then what you have to select is the

Python environment. And luckily if you

installed Anaconda

you have a built-in Python environment

which is going to be your uh which is

going to be um the

you know which is going to be the uh

Anaconda that you installed.

So I have Anaconda here. Now I have a

lot of other ones but the

uh Anaconda is here. Say it's this one.

Does that so when you

when you uh

Yeah, you have to install you have to

install a Python environment. Yes. So

you want to install Anaconda first and

then you can run your then you can run

your notebooks.

And then you just uh run your code as

usual

and then we can run that.

Yeah, it but like it's working as if you

know the same kind of notebook that we

have inside of Collab, the same kind of

notebook we have in Jupiter.

You you have to have a Python installed

for this to work. So you go to you go to

Python environments

and then choose a Python environment.

You could try to create one. I'm not

sure if that'll work for you. Create

Python environment. You could try that,

too.

Yeah, that's fine. Any anyone will work.

Any Python will work. You just need to

pick a Python. Anyone will work.

Okay. Yeah, if it defaulted to something

that's fine, too.

And then we can generate more cells

and run cells.

But yeah, that's the thing is you're

going to want to install um Anaconda

most likely because you need a Python

version installed on your machine in

order to run this.

Perfect.

Okay.

All right. So, what I want to do is jump

back over to our notes so we can

continue along. Um again feel free to

use whatever

platform works for you. Collab,

doing notebooks in VS Code, doing

Jupiter notebooks, whatever works for

you, please feel free to use that. There

is no wrong way of using it. Whatever is

best suited to you and you're most

comfortable with, please use that

option.

Uh, it's lowercase P. That's why

lowercase P. Capital P is not a function

in Python.

Lowerase.

Yeah, go with Collab. Yeah, if you're if

you're on a machine, you can't install

anything, go with Collab. That's totally

fine. That's why it's there is for the,

you know, convenient kind of cloud

aspect to it.

Um,

feel free to do collab for everything.

That's totally fine.

I will use collab from time to time as

well.

All right, let me uh go back to our

notes then

and pick up from uh syntax. So, I'm

going to go back to Let me share my

screen. Go back to

Can you use Collab on your phone? I've

never tried it. I would be surprised.

Maybe an iPad.

Maybe like a tablet. It could work

pretty well.

Phone. I'm not so sure.

Yeah. Go ahead. Try it and let me know

how it works.

Try it and let me know. I really don't

know. I'm curious now to try that.

All right. So, I'm going back over the

notes. We're going to finish up today uh

what the time we have left to go through

some syntax. So really uh

really getting into um into Python like

the actual code of it so we can get

started on that and start working our

way through it.

Uh the difference so the the difference

is um you will be executing py files

with the within the terminal. So you'll

be running Python files instead of cells

in a notebook. you're just you're

running a you're running a Python script

rather than individual cells.

Okay, so there's a difference there.

And the reason the reason we choose

notebooks is to run individual cells.

It's just easier.

Same syntax,

same syntax. It's just the code is not

isolated into cells.

still Python.

All right.

Um, let's go forward into the syntax,

start learning about it.

All right. So, something we need to

learn about is how do we properly write

Python code? What is the syntax to it?

So, some things we're going to need to

learn about are how to write proper

identifiers, which are names.

Identifiers are just names for

variables. So, we need to know what's

allowed, what's not allowed. We need to

talk about what the indentation means

and why do we need it in Python. I want

to show you guys how to write comments

because that's really important to

leaving notes to yourself or others

about the code and then talk about um

generally how we produce output and how

we can accept input um from a user or

someone interacting with our code. I

want to talk about all these things.

We'll see how how much of what we get

to.

But let's start with the identifier. So

what is an identifier in programming?

This is really for any programming

language. An identifier is just a name

we give to something inside of our code.

So it's a name we give to a variable, a

name we give to a function, a name we

give to an object. Um so any name we

give to something in our code, like when

we set something equal to x, like x

equals 3 + 3. um that thing the x is the

name we are giving or assigning to a

result or a variable or an object. So

anytime we write down a name in our code

of something there are certain rules

that those names have to follow and

these are something we will um pick up

as we go along but I wanted to call them

out here. So um when we name anything in

Python

generally they have to follow these set

of rules meaning they have to be a combo

of lowercase or uppercase letters either

one's allowed it can be even be a

mixture of lowercase and uppercase

numerical digits are allowed in the name

that's okay any digit 0 to nine it's

okay and underscores are okay

underscores are Okay. And there's no uh

minimum or maximum length. So names can

be really long, they can be really

short. Um of course they should be

meaningful. So when we name something,

it should not remember we want to kind

of get away from naming everything X or

Y or A or B um because those names may

not mean much when we look back at the

code. So even though those are valid

names from an identifier perspective, we

want to be really meaningful when we

name something. We name a variable, name

a function, name an object. Um,

here's one catch.

The name cannot start with a number. So

I can't name something uh just the the

number zero or the number one um because

I can't start with that. Now, it can

include that.

So, if I need to include a number in the

name, as long as it doesn't start with

it, that's okay. But names cannot start

with a digit. That's just one rule of

Python. Anything that you're assigning a

name to,

like a variable, function, whatever,

cannot start with a number or else it'll

be invalid.

Okay?

So I'll show us examples of that later.

Yes, they can start with underscores.

Yes, that's okay. It can start with

underscores. Of course, it can be lower,

uppercase. It can start with It cannot

start with a digit. It can have digits.

They just can't be the first character

of the name.

Yes.

Um, now special symbols cannot be used

in the name. So you cannot have a

percentage, dollar sign, exclamation

point, hyphen,

pound symbol, at amperand symbol, at

symbol. None of those can be used in the

name. So those symbols are not

recognized.

So if you try to include those in the

name of something,

that will produce an error. So we don't

want to do that. The other thing we want

to avoid is naming something in the same

name as something that already exists

internal to Python. So those things are

called keywords. So there are certain

keywords that have a meaning in Python.

They are built into the language. We

cannot reuse those. They're basically

reserved. Um so something like class is

reserved because it means something. It

means you're declaring a class. We'll

see. We'll talk about what that means

later. Or something like global cannot

be used because it declares something as

global. Um,

so

you know, we'll learn what some of those

keywords are. There's a list of them

that are in the Python documentation,

but we want to avoid naming things after

builtin

uh uh functions or builtin keywords. Um

so so in fact we've already used one

which is the print function. You know

when we printed out hello world we would

want to avoid naming something print

because print means something. It exists

as a function. We don't want to name

something print

right that would that would produce it

because it would produce confusion. The

interpreter would see that and say oh do

you mean the function print or you

trying to name something print? it

wouldn't know. So, we want to avoid

naming something that already exists

inside of Python like print or like

class global. Um, there's many others.

Okay. Lastly, and this one always throws

people off, is that when we name

something uh that is case sensitive. So,

if you name something lowercase A, that

is a completely different variable or

completely different function than if we

were to name something capital A. These

are different. They're treated

differently. So, Python will think that

those are two different uh objects or

variables or whatever the case is. So be

really careful with case sensitivity.

Python is case sensitive.

Lowercase A will not be treated the same

as capital A. And whenever you're naming

something, so if I have a variable and I

I I set lowercase A equal to five and

then I set um uh capital A equals to 10.

Then if I um print A, that would produce

five.

But if I print capital A, that will

produce 10. It's not the same. So it is

case sensitive.

These would be two different names.

Lowerase A and capital A.

Okay.

So, we're going to do examples with

these, but these are just some rules we

have to abide by in the syntax if we're

naming anything like naming our

variables, naming our functions, naming

our objects as we go along in in the

course, right? We just cannot The main

one that trips people up is we can't

start with a digit and we can't use

words that already exist like print.

Oh, is my video stuck for people?

Was it stuck?

Oh, okay. Always let me know because it

might it might be.

Always let me know because it definitely

could be. So, it's better to know than

not to know.

Okay.

Oh, no worries. Like I said, always

always feel free to to let us know cuz

um it would if it is, then it's good to

call it out. So, no worries about that.

Any questions about these names? Do does

it make sense about like how we name

things matters and there just are

certain rules that we have to follow. Um

we want to avoid these wacky symbols.

Um, you know, we want to avoid naming

things that already exist. We want to

avoid starting with a digit. Otherwise,

it's going to be a pretty standard like

lowercase, uppercase mixture,

maybe occasionally with an underscore

mixed in there. Um, or or digits even.

As long as we don't start with one,

that's okay.

Yeah, Tim, that's a good reference. the

PEP. So PEP

are the set of guidelines that um the

Python Foundation has kind of agreed

upon as um here's what you should use as

your style guide. Here's what here's

what the community believes is the best

style for Python. Those are good to

read.

Yeah, those are those are uh good

references for really like formatting

and styling your your Python code uh to

be in line with kind of what the

community expects.

Okay.

Okay. So, let me give you some examples.

Um so, the ones on the left are going to

be valid. So, we can name something my

class. We can name something var_1

that's okay.

Count

um that's okay. Uh

but if we have

um like on the right if we have

something that starts with a digit that

would be bad. So so this one is no good

because it starts with this number.

That's not good. Um this name has this

wacky at symbol in it. that's no good.

So, this would be a bad name for

something that would produce an error.

Remember, the interpreter is going to

see that and reject it essentially and

say, "You can't name something this.

It's not valid." Um, same thing with

trying to name something global. This is

a keyword that already exists in the

language. The interpreter is going to

see that and get confused. It's not

going to know if you're talking about

the keyword that's built in or you're

trying to come up with your own name.

It's not going to know. So, it's just

going to throw an error.

Um, so again, we want to we want to keep

things simple. We want to use

underscores where it makes sense. We we

don't want to start with numbers. Um, we

can use a mixture of lowercase and

uppercase. That's fine.

Um, this is a good variable name rather

than if I just called something X.

That again, we're trying to avoid that's

something I always see in the beginning.

I think is okay in the beginning, but

it's something we really want to be

conscious of is naming our variables

very meaningfully.

Like count is going to be more

meaningful if we're keeping track of a

count of something. We would rather call

that count than if than if I just called

it X. Because if you read the code,

which do you guys believe me? Like when

you see it, you kind of know exactly

what it means. X or count? What do we

think?

Which one like has more meaning to it

when you see it? You know exactly what

it's keeping track of. X or count?

Yeah, count.

I would agree with that. Count. Yep.

So, it's just an like that's just a

single example of trying to keep track

of things in a meaningful way. That's a

good name to give to a variable. That

would be uh you know keeping track of

something the count of something

rather than if we just called it x or y

or a or b.

All right, I want to talk to you guys

about indentation next. So now we know

we have to name things appropriately and

the interpreter will give us an error if

we don't name things appropriately.

What about indentation?

So indentation

is a way for Python to understand what

code gets executed together.

Okay.

So

and it it also indicates that I am

breaking the flow of the code from one

section to the next. So the indentation

is really important to signify to the

interpreter there is a new section of

code that has to be considered

um before I move on. So um you should

always use indentation

whenever you have a colon like we see a

colon here with if else statements. Now

we haven't learned about if else

statements but we will. But notice how

we have the colon there and the

interpreter would be okay with this.

This would work.

Okay,

which is a simple statement of saying is

five greater than two? Yes. So in the

case that it is, let's run this code.

But we're only able to run it because

the interpreter recognizes it's

indented.

So the indentation is really really

critical as it makes the interpreter

understand what should be next. The

interpreter understands what should be

next after this statement. Um like an if

statement or a loop statement. Um we

will always have indentation. So this

would actually uh throw an error because

there's this is not indented. This is at

the same level and if you have collab

open you could try this for yourself.

Um if you had collab open it you could

try it for yourself is like this would

this would throw an error where it says

I am expecting indentation but you did

not have have any.

Does it matter how many spaces?

Uh you yes you want to use four spaces.

This this should be four spaces here.

Spaces or tabs?

Uh I'm only laughing because it's a

it's a kind of a controversial question

in the community. Some people get really

upset over

one or the other. I'm not one of those

people. I don't really care. they so

most code editors

uh set the tab automatically as four

spaces. So a tab will do the same thing

as if you manually just did four spaces.

It doesn't really matter in that case.

So either way,

yeah, you're so the ide will do that for

you. The IDs will generally do that for

you because they know it should go on

the same line.

Would it work as on the same line? Yes,

in some cases it will, but not all. It

depends on how complex it is. But what

do you think is easier to read

from a readability perspective? Which is

easier

if it's all in one line or is it more

readable and easier to follow if it's

indented?

Yeah, the that's the purpose. So yes, I

I agree. indented makes it easier to

read. So that's another reason Python

really enforces indentation is to make

it easier to read. There's a reason they

do that. It's to make it easier to read.

Okay?

And that's one of the best selling

points of Python is how easy it is to

read and work with. The indentation

really helps. So to summarize this, we

are going to have to use indentation.

Anytime we have

uh a statement with a colon. Anytime we

have a statement with a colon, we're

going to have to have an indentation

immediately follow it. And there are

certain statements that have a colon

like if, else, else if, and any loop,

any looping statement. Now, all of these

things we're going to learn about, we'll

learn about it in our next lesson.

But anytime we have a colon, this is

signaling the interpreter, okay, there

needs to be a block of code following

that, which is, yeah, as you say,

Romero, it's like a a hierarchy. Yes,

that's a great way of thinking about it.

It's saying, okay, I should check this,

then do this if that's true.

that tells the interpreter this is only

going to be executed in the event that

this is actually true. Otherwise, I'm

going to keep going.

Okay.

All right. Any questions about the

indentation? This is something we're

going to learn about more as we go

along. When do we use indentation and

when do we not? We're going to learn

about it when we get into the if else

and the loops which we will study.

But do we do we let me ask you guys

this. Do we understand the idea or the

intent behind indentation?

Do we roughly get that idea? We don't we

don't know yet when to use it. I get

that. But more the intent or the purpose

of using it is to really like section

things off.

Yeah.

Okay. Good. Good. Good. Good. Glad to

hear that. Okay.

Okay.

Let's wrap up today by talking about

comments. So, uh what are comments?

These are like annotations or notes to

yourself that are completely ignored by

the interpreter.

So when your code gets executed, the

comment does not play any role in what

gets executed. The interpreter will

actually just completely ignore it. The

moment it sees the comment, it will just

ignore it and go to the next line.

So the purpose of it is for humans to

leave a note to other humans reading the

code and that is very powerful is to be

able to read those to to leave those

notes and not have it affect the actual

code that's uh that's actually executed.

So there's multiple ways to make

comments inside of Python. The most

basic is to use the pound symbol. So the

remember we cannot use pound symbols to

name anything.

This is why because the pound symbol is

res reserved for making comments. So you

you when you have a pound symbol like

this uh that immediately signals to the

interpreter everything else that follows

that on this line is a comment. Any

other text that follows that on that

line is a comment. And usually your

editor like VS Code, Jupiter, Collab

will color that differently, maybe like

a a like you can see in here, this is a

Jupiter example. You can see it's kind

of a light gray,

greenish gray

um to signal that this is a comment. Um

now you may be asking why would we have

comments? Again, you are going to look

back at code weeks later,

especially in this program. You're going

to look at code in review and be like,

what what was I doing there? If you

leave a comment, you'll remember what

you were doing there, why you had that

line. Um, and not only that, like when

you share your code with others, which

in the real world you would be doing,

collaborating with others, right? Adding

in those comments can be really

beneficial to do. So

you will see me

throughout the program. I'm going to

leave a lot of comments on our demos and

our notebooks that we work on together

in the live sessions. I will leave

comments mainly to call out certain

things like I will say this is a really

important step or we are doing this

because I will leave a lot of comments

and I encourage you guys to do the same

in your own code. Um, remember they're

free. They're they get ignored by the

interpreter. They don't affect anything.

They are notes to yourself. So, use them

accordingly. Um, and you know, there's

actually multiple ways to leave

comments, but this is I'll show us those

as we go along. But this is the uh most

basic is you just you you type in a

pound symbol and then everything else

that follows that is uh is your comment.

Okay.

All right, guys. That's it for today.

Um, what a great first session. Thank

you guys. Thank you guys for being

patient. um going through the setup of

some of those tools. I hope you landed

on one that worked for you. Um you know,

use that one going forward, please. If

it's collab, use that. Jupiter, use

that. VS Code, use that. Whatever you

you uh feel most comfortable with,

please use that. Um we have a lot to

cover still. You know, we're going to um

continue on Wednesday. Uh we were we

will uh continue talking about Python.

We're just getting our feet wet a little

bit on on Python. A lot more to cover.

We're going to get into the the more

nitty-gritty of the code. So, it'll be

really fun. We'll cover if else loops,

um how to control the flow of our

program. Um we will do that. This is

where we left off was writing comments

uh in Python code, which I I did want to

remind us of how to do that. it is going

to be using the uh pound symbol to

initiate a comment. And basically the

Python interpreter will ignore

everything else on that line. Uh it it

treats all of that text as a comment.

And again like comments are free. You

might as well use them to your advantage

to kind of uh leave a note to yourself

of hey this is what this code is doing.

Um so that when you come back and read

it uh you can understand it better. So I

encourage you guys like when we do demos

uh and we will do a lot of demos um

especially today leave comments you know

put comments in there so so you can make

a note to yourself what this code is

doing um so you will get I think it'll

be good to get in the habit of leaving

comments uh to kind of mark up the code

to to kind of remind yourself oh this is

what it was doing when you look back at

uh in the future.

Okay. So we had ended with that.

What I wanted to do was move on into the

next slide. So talk about um basically

how we display output to the screen

which we've already seen an example of

when we did the hello world which is the

print function on the right. So this is

a by the way this is a Python function

and you know it's a function because

of these parentheses. So these

parenthesis signal that this is a

function because it expects some sort of

input to go inside of those parentheses.

And and the input that would go inside

of there is going to be text like some

sort of uh some sort of text that

belongs inside of quotes. and whatever

we put there um will display on the

screen. So that's useful for us to like

display information

um print we would say we are printing

out information to the screen. Um so if

we want to know the value of a variable

or the value of something that we are

doing a calculation with or uh you know

sanity check something in our code we we

can print it out which would be using

the print function and it will display

that value onto the screen. So we will

use the print function quite a bit. Um

you know we haven't learned what

functions are but uh functions in Python

are you know um designed to be uh chunks

of code that execute and do something

and they take arguments and you know it

takes an argument because of the

parenthesis that is um signaling that

there should be some something inside of

this parenthesis here which is going to

be uh text. So whatever text you want to

display or maybe some variable you want

to display um that would go inside of

there. So we'll get the hang of using

the print function as we go along but

just wanted to call out that's the

primary methodology of kind of um

displaying something on the screen if we

want to print function. Um now the

reverse of that is uh asking a user to

uh input some data. So that would be

this input function and um this is

something that uh as you can see an

example below is we can put some text

inside of this parenthesis. So again a

function it has those parenthesis that

signals it's a it's a function. Um,

and we can put some text in there which

would be kind of what displays in to the

user as kind of a prompt like here. Uh,

enter your name and that would display

on the screen and then there would be a

box next to it. I'm going to show us

this. I'm going to actually run this

inside of a notebook in a minute. But

then there would be a box that displays

that that would say um hey you know

enter your name and then you can type

input in uh and and then when you hit

enter it will save that input into into

this variable called name. So um and

remember name this is a valid identifier

because it starts with a lowercase n um

which is fine and it it has all valid

characters. It doesn't have any wacky,

you know, uh, pound symbol or at or

anything crazy. So, it's it's a decent

identifier. Um,

so name name would be okay. And so input

is whenever you want to get whenever you

want to allow the user to input

something like it'll bring up a text box

and they can enter some data. Um, and

that will be saved in this whatever

variable you name this you set equal to

input. Um, and then you can see like as

soon as we put that in, we immediately

dis we can display it. So we we print

hello and then comma name which

references whatever we stored whatever

the user input there. So I'm I'm going

to show us an example of that. Um, but

input is what get is our primary way of

getting input from the from the the user

in a text box so we can use that data in

our program.

Print is our primary way of displaying

data that we already have in our code.

We can print it which will display it.

Um

so we're going to see many examples of

these as along but just wanted to call

out those two. These

functions by the way are built into

Python. So we don't need to create them

ourselves. They already exist. They're

already built into Python. Um nothing

special we need to do to use them. We

can just use them right out of the box.

So again, we'll see this in our in our

code examples that we're going to do in

a minute.

Um, where exactly would an end user be?

So maybe we ask them for some input. Um,

and then we do like so we ask them for

some like their name, their email, their

uh date of birth, those kind of things.

We can ask in the input and then we

maybe we store them in a database or we

do something with it in the Python code.

So um whenever you want to accept input

from a end user that's when you would

use this input.

It just depends on the application

right on the application like what kind

of input data you you want to accept

from the from the user.

Okay. Okay. So, I'm going to show us

this.

Um, before we do that demo,

um, let me ask you guys, which of the

following do do we remember from Monday?

Which of the following identifier names

follows Python's rules

and best practices for readability?

So not only so you should be looking for

the answer choice here that follows the

rules but also is a meaningful name.

A lot of different choices. Okay.

By the way,

which let me ask let me I'll come back

to the answer to the original question,

but let me ask this alternative

question. Which one of these is not

valid? Meaning it would Python would

throw an error how you use it.

Which one of these is not valid in

general?

Cool. Great. It is C. You guys were

right on top of that. Very good. So C is

not valid. It names of things cannot

start with a number. So So that's um C

is completely invalid in general and

that would produce uh an error.

Right. Starts with a number. Exactly.

Which we cannot do. Now starting with an

underscore is okay. that's allowed. So

that's not an issue. And having a number

be second after the underscore is okay

as well. So technically A and B would

follow the rules. So those definitely

follow the rules. Um now are they

readable and meaningful is the question.

I would argue that possibly not. Var 123

is pretty generic. I would argue that

even though it's valid, like Python

would not have any errors with that uh

variable name, it's not very meaningful.

It's too generic. It's almost as if we

just called something X. We just called

something var 23, that's probably not

going to be meaningful to us and we're

not going to understand what that really

represents. If somebody were to come

along and read it and see VAR 123,

that's probably not that great of a

name. It's not telling us exactly what

that represents. So, I would say A is

likely not um

A is likely not uh a good choice and C

we know is invalid. So, really I think

the only two options you could argue are

B and D. I think D is a really good

answer. It um it it follows the rules.

Uh underscores are fine. Everything is

lowercase. That's fine. Um so it's valid

but it also is meaningful as a name

right so final result value um we we

should probably like in our code we

would have context we would know what

that means okay this is our final result

um so that's a that's a pretty good name

for something um you know this one is

okay it's just not that readable 321

customer details DB table it's okay I

don't think It's um it's not the worst.

It it definitely would work, but it's um

kind of a clunky name. I'm sure we could

come up with something better, but it

would work technically. There'd be no

issues with it.

Okay. So, I think D is probably the best

choice, but B is valid, too. I think B

could work for this. D and B, I think,

are okay.

Okay. Good. Good. You guys are right on

top of that. you have a good I think you

have a good feel for what the allowed

names for things are which is good.

Okay.

Um

I'm going to then swap over to this

demo. So you guys should have the uh

demos um and we kind of went through

some of those first few last time to get

you set up on Cola and Jupiter and VS

Code. Um, I am going to be using Collab

for most of these, but feel free to use

whatever you want to use. If you want to

use Jupiter, if you want to use VS Code

and run your notebooks on your own

machine, feel free to use whatever you

want to use. I'm going to be using

Collab just for the simplicity of it.

Um, so, so this demo will walk us

through, um, opening up a new Collab

notebook and then running those input

and print. So, some examples with input

and print. Um, so we'll do that

together. Let me go over to that demo.

So, if you're following along, we are

going to be doing um demo 4. So, it

should be lesson one, demo 4.

Um, do you guys have this? Give you a

moment to to pull that up. Lesson one,

demo 4.

You guys have access to this one. So, we

did we did one, two, and three on

Monday, which were just getting those

environments set up. So, this is demo 4.

Um, which again, I know step one says

open collab. Feel free to open your own

notebook in VS Code or open your own

notebook in Jupiter as well, whatever

you're most comfortable with. Um,

I'm going to be using the collab to to

do this, but feel free to use whatever

works. You're we really just need a

notebook to be able to run this code.

So, however you're running notebooks,

whether that's in VS Code or Jupiter or

Collab, I any of those, either one is uh

perfectly fine.

So, so step one is to open up a

notebook. I'm going to do it in Collab,

which is what this says. Um, and then

you can make a new notebook and then

rename it to my first program. I'm going

to do that in a second. And then, um, so

I'm going to walk through this live with

you, but just showing you some of the

steps we're going to do. Um, the first

thing we're going to do is is just

practice doing the print hello world

again so that we can um, execute a print

statement. So, we'll practice that.

We're going to make a second cell. Um,

which we can do in Collab or VS Code or

Jupiter by hitting the plus button.

There's usually a plus. Uh, you can see

it here. Uh, in multiple places in

Collab, you can do it right below an

existing cell or there's always a plus

code here, which is kind of what you

have in Jupiter. Usually in Jupiter, you

have a plus button. So, you can just hit

hit that plus button, it'll make a new

cell. Um, and so we'll make a new cell

so we can write some more code.

Um,

and then in this new one, we are going

to practice doing some comments.

We're going to practice doing some

comments and then um see how we can do

uh some more print statements. Okay, so

let's do that. Let me jump over to

Collab. Let's walk through these first

few steps together. Um, and then uh

we'll come back to this and finish out

the rest of the steps because we're also

going to do input. So I'm going to show

you how to do these uh input which will

you can see here like the input is going

to create a text box where you can put

input and it will you hit enter it will

save it for you. So input allows you to

get input from the keyboard

and save that into a variable to use for

later.

Okay. So, let's jump over to

um let's jump over to

I'll show us I'll show us in a second.

How do you rename it?

I'll show you. Let me jump over to

collab.

Um

okay.

So, I am over lesson one, demo 4. Yep,

that's the one we're doing.

Okay. So, I am in Collab. I'm going to

start a new notebook.

Start a new notebook in Collab. Uh, so

now I'm here. I'm just on a fresh

notebook. Um, nothing that interesting

going on. Here's how you rename it. just

go up to this box on the left

and almost like a Google doc just just

uh click into that name and then start

typing to erase it. So see how I'm like

hovering over that name and then I'm

clicking on it and then I can start

typing to erase it. So I can we can name

this my

first program

and then hit enter and it will save

that.

Oh yeah. So if you're in VS Code um do

to to rename it do file and then save as

and then you can give a new name to it.

file, save as.

Okay, that's how you can rename it in VS

Code.

All right, let's do let's do the first

step. Um, let's do print. So, we're

going to do print. So, type in print and

we can uh we can do parenthesis.

Um, remember this is a function. So, we

need the we need the parenthesis to

signal that we want to put some text

inside of this print function. And then

you want to do uh you want to do quotes.

You want to do quotes, the double quotes

there, in order to allow us to put in

some text. So, Python will interpret

what's inside of the quotes as text and

it will display that text. So we can do

hello world my first

Python program.

Okay. And then we can run it. So feel

free to put whatever text in here. It

doesn't really matter exactly what it

is, but you put some text in there

between the parenthesis and then hit

run.

And the notebook will take a second to

connect. And then there it is. Right? So

then you see the the text displayed on

the screen.

Try that out. Are you guys able to run

the print

in your Jupyter what whether it's

Collab, whether it's uh Jupyter

notebook, whether it's VS Code. Can you

run the print

install?

Yeah, I installed that um in VS Code.

Yep. Try installing that.

Okay, great. You guys were able to run

that. Very good. Very good. Okay.

No, you don't want to save it as a JSON

file. You want to save it as a pyb just

like this. See how this one is uh IP

YMBB?

That's the format you want. Remember

that is interactive Python notebook.

You want that file IPY MB.

Uh perfect. Yeah, you get you got it to

run.

You don't have any extension? No. If

you're in VS Code, remember from Monday,

you need to install the the Jupiter

extension.

If you're in VS Code, you got to install

the Jupiter extension.

You have to manually so manually save

it.

You can save it as a py. I would do ipy

so you can open it in collab.

Type it yourself.

Type overwrite what's there and type it.

Type in um my notebook whatever the name

is.

Type it out yourself if you can. like

save as and then

type out the full file name yourself.

Now let's practice a comment. Let's

practice a comment. So let's build let's

do a new code cell. So we made a you can

either do it here. If you hover over

your cell, you can hit plus to build a

new code cell or you can hit plus here

to make a new code cell. So let's do

that.

You should be saving.

Don't worry about the type. Just type in

the namey imm. I don't think you need

to.

Or you could just hit what you could do

is you could just hit save and then in

your file explorer you could just rename

it.

So if you just save it will save it to

the default location and then just and

then just rename it.

So maybe try that route. Just just do

save. Just save it. And then it should

it should try to save it as IPymbb.

Okay. Okay. Let's practice. Um, so the

next step in the demo, if you're

following along in the demo document, it

wants us to do, so we did the print. We

want to do um a practice some comments.

>> Okay, perfect. Uh, let's practice some

comments. So, um, remember I told you

that we can do, uh, comments with the

pound sum. So, this is a

comment. So practice writing a comment.

Remember you start a comment with a

pound symbol.

Um it will get ignored

by the interpreter

interpreter. So feel free to type in

whatever text you want. I'm just

reminding us that whatever the comment

is is going to be ignored and we can

have whatever code below that that we

want to have and that comment will get

completely ignored. So let's do another

print. So write a comment,

hit enter. Immediately below that in a

new line, let's do another print.

This code

gets executed.

So we know this print statement is going

to get executed, but this comment is

going to be ignored by the interpreter.

So let's run that.

So this code gets executed. This comment

gets completely ignored,

right? That comment gets completely

ignored, which is great. Try writing a

comment. Are you guys able to write

comments?

So write a comment and then try writing

a print statement right after it.

And and feel free to put whatever text

you want inside the comment. And feel

free to

uh for the comment, is the space after

the pound symbol required? No, it's not.

So, we could test it out. So, I removed

the space. Doesn't matter. It It's just

for readability. I usually like doing

that so that I have some space after it.

And this is a little It's just a little

bit more readable, right? It's not like

mixed together.

It's just for readability.

Great. You guys wrote a comment. Okay.

Perfect. Perfect. We're able to write

comments. Really great. Okay.

Okay.

If you put multiple code lines, do we

need any separator? Like, no, they just

go on new lines. So, do you mean like a

second print statement? Let's We could

try that. Let's do a secondary print

statement. So, we can do print.

Um, this one is on the next line. No

separator

needed.

Do you see that? See how it's on its

own? I did a print right below this

other print. And as long as they're on

their own line, that's okay. They just

need to be on their own lines. They

don't need any separator.

If we run this, then this one gets exe.

Then see how this is now printed out

below it. Right there.

Is there any character limit on the

comments? Uh, no. There's no character

limit. Um,

but

there's no character limit, but a good

practice is to not like you don't want

this to be super long and to to take up

the whole screen, right? Because then

it's not really readable.

So, there's no limit, but you don't want

to have overly

long comments. You want to keep them

kind of concise and short.

So just so you can read them and they're

they don't take up a lot of space.

Not able to add print statement below.

Why? Why is that?

You should be able to should be able to

have a print right below this print.

Shouldn't be anything that make sure you

close this parenthesis. Make sure every

print needs to close the parenthesis

and they all you also need to close the

quotes. So close this quote, close this

quote within within the print

that needs to be done. So you should be

able to run I'll paste this for you guys

in the chat. Should be able to run this

All right. One thing I want to show you

guys is just like the demo says in the

word document, um you can do multi-line

comments. So if you need to do a lot of

comments, all you need to do is triple

quotes. So triple quote,

then um triple quote, and then

everything in between.

That's interesting that it did that.

Yeah. So, we can do a pound symbol,

pound symbol, pound symbol,

pound symbol, and that that should all

work. So, we can do that.

Yeah, I think it's a collab thing, but

normally in in like Jupiter or in

Python, it it will work just fine. But

like in collab, I think they don't like

the triple quotes.

But yeah, do you guys see do you guys

see how I just did it like this with the

pound symbols? That's okay, too.

Everything between these

uh pound symbols is a comment and is

ignored. So now we should be able to run

that. So there we go. Everything gets

ignored there. Does that make sense to

us? The pound symbol comments

does the does the using the pound

symbols. So notice how we use that to do

multiple lines of comments. So we did

one here, we did one here. We can have

as many we can have

um as many

uh comment lines as we want and they

will all get ignored.

What is those? It's supposed to be

multi-line comments, but for some reason

it's not working. Um, it so the the

triple quote is supposed to be like

representing that you can have a whole

block of comments.

I don't know why it's not working in

collab for me.

It's working for you. Okay. Okay. I

don't know why it's not working.

Single quote.

It's still It still displays here, which

I don't get why that's happening.

It's kind of weird to me.

Yeah, I don't get why inconsistency. It

usually It usually works for me. I don't

get that at all.

Still still doesn't work for me. I don't

know why that doesn't

Yeah.

I don't get why that's not really liking

those triple quotes. Oh well. I mean,

not a big deal. We can just do

Okay, we can do we can try single.

Still doesn't work.

Yeah. Oh well, we can do a pound symbol.

That will always work. Pound symbol is

honestly more popular anyway. Most code

that you see in the wild will have pound

symbols wherever they're doing um

wherever they're doing uh comments. So

that it's fine. Just use a paddle for

now.

Uh yeah, that's correct. I don't know

why that's that's correct. Um I don't

know why collab doesn't seem to like

that. It should be ignored

um generally with the triple quotes, but

uh that's okay. I'm not too concerned

about it for now. I guess what you and

when I do comments, you're usually going

to see me using the pound symbol

anyways. It' be very rare that I would

need to do uh quotes.

Yeah, it's weird that collab doesn't

work very consistently. That's okay.

All right. What I want to show us is I

want to move on to the input. So, I want

to I want you guys to see

I want you guys to type in this code

here that will take input from a text

box and save it into a variable called

name. So, the code we're going to do is

going to be like this. It's going to be

name equals input

and then we'll put um please

enter your name.

Okay.

So this, by the way, I'm going to

comment this code here. Um, this code

should

create

a text box for us to put in our name.

Okay, so that's what should happen. So

when we run this, um, it should pop open

a text box right below this. And we can

type in our name and hit enter. And when

we do that, it will store that result in

this variable called name, which we can

use uh wherever we want to in the code.

So if I hit run, there's that text box.

Do you guys see that? There's the text

box. And see how it says, please enter

your name. And so we can type in our

name.

And we hit enter. And there it's stored

in the name. We can even um display name

by doing print and then the name which

will display uh the name that we stored

when we did the input.

So try this one out. Try this code out

for yourself. Try typing input

parenthesis

and then you want to have some text

there. It doesn't matter exactly what it

is, but something like please enter your

name or enter your name.

Try that out. And then it should store

uh you should be able to type in the box

that shows up. Hit enter on your

keyboard. It should save that. And then

you can um print it out. You can print

out that name which will um

display that whatever we typed in

before.

What does it look like, Roberto? What

does it look like? Were

you Were other people able to run this?

Oh, yeah. Thank you, Melanie. Yeah, I

see that. Perfect.

Perfect. That looks good to me.

Uh, you don't need a space. um it just

looks nice, right? It's so that's a good

practice to have the space so that uh

this code is um evenly spaced out and it

looks nicer on the on the screen.

Um name equals input print hello there

uh name

You need Yeah. So, uh, Roberto, you need

a you need a comma after after the

quotes.

After the quotes, you need a comma after

the quotes to signal to Python that

you're putting in you have you have some

text and then an additional input.

So, it need it needs to be more like it

needs to be like this. print. Um, hello

there.

And then you need an extra comma.

See how I have an extra comma after the

quote. You need you need that.

Sorry. Now, Roberto's uh Kiati.

Hope I'm pronouncing that right.

Okay. So, do we feel good about input

and what it does?

Perfect. Do we feel good about input and

what it does? It It brings up a text

box.

It Did you hit enter, Roberto? To like

Were you able to type something in and

hit It's going to run until you hit

enter.

You have to type in the text and then

hit enter into the box.

So, let me rerun this. So, it See how

it's still running? See how this like

it's going to keep running forever until

I type something in

and then when I hit enter it will stop.

What does your code look like?

Okay, that looks right.

Try try stopping it and rerunning it.

Try try hitting the stop button and then

rerun it.

um you should so yeah you should put

that in a different cell. So if you if

you separate your code into individual

cells so you could do like um you could

do name. So we could we could separate

this. So this code is the only code

that's running in this cell.

That doesn't make sense. Something else

is

that doesn't make sense cuz like this

collab tab is only taking up 235

megabytes.

So something is

chewing up your memory that's not really

I I can't imagine. Are you using collab?

You can see like it's not using that

much. Only 240

230ish.

Yeah, I don't think I don't think Collab

is the culprit unless you loaded in some

really massive data or something.

I can't imagine that's the issue.

You did. You loaded in data. That's

I mean Yeah. Then it's going to it's

going to take in memory. Oh, okay. Okay.

Okay.

Uh MJ, what are you on? Are you on

Yeah. Could you screenshot it?

If it's not working for you, could you

try collab? If Could you try collab just

for the sake of like getting it running?

Things should work in Collab pretty

easily.

You're using Collab and nothing's

working. Uh, are you making sure it's a

code cell and not a text cell?

It's not a text cell like this,

which would be like,

this is where it will be blue.

Did you have that? You need to make sure

it's code. Yeah.

And when I run that, it's going to be

it's going to display text. Yeah.

Okay. Great.

Glad that it's working. Great.

Okay.

Um All right. One more. Uh one more

example what I want to show you guys is

how to do how to include the name in a

print statement. So if we do something

like print. So um we can include the

name in a print statement. So if we do

something like print and then we have um

hello there and then we have um this and

then we have welcome to Python.

um this will

uh this will display all of that

together. So notice that we can have as

many um pieces of information that we

want to display kind of one after the

other as long as they're separated by

these commas.

So we have this uh text,

this text because text is stored in that

variable. Um, then this text and then we

print that all out and we can have this

whole collection of text displayed to

the screen. Try that one out.

Oh, they do the same thing. They do the

same thing. So, the comma and the Sorry.

Yeah, I just noticed the demo does a

plus. They do the same thing in Python.

So, we can swap that over to a plus.

Both of them work.

They have the same I shouldn't say they

do the same thing, but they have the

same effect.

They have the same effect.

Actually, there's no You need a little

bit more spacing here. So, the comma

gives you a little bit better uh

spacing.

So, what the Let me break this down.

what the so plus

plus um adds together

uh text and so what we're doing here

technically is adding all our text

together and then displaying it. Um, so

plus as together text and then the comma

um,

uh, prints out multiple pieces of text.

So they they have the same effect, but

yeah, you can use you can use either

one.

Okay.

Um, one thing I wanted to show you guys

too, by the way, in Collab, if you're

working inside of Collab, I want you to

hover over your name variable.

So, if you just take your mouse and

hover over that,

do you guys see what it says here?

Do you see how it says string name and

then it has the value of that, which is

which is my name. So, that's something

cool about Collab is if you hover over

variables, it will tell you what their

type is. Now, we haven't learned about

types, but any text inside of quotes is

a string. It's it's what we would call a

string. We're going to learn about that.

And

um

we it also displays what data we

currently have stored in that variable.

So all you have to do is hover over a

variable um to to see what the value is.

Yeah, that's yeah, that's kind of a

limitation of VS Code. That's true.

It doesn't show you immediately on

hovering.

Don't see the value on hovering. So, um,

click into the cell. You have to click

into the cell and then hover over it.

Click into the cell and then hover over

it. It should it should work. Yeah, you

have to click on the cell or whatever

cell you're on and then uh hover over

that and it should work.

Uh, Mariel asks, "How do we integrate

that Python code to a client application

for a user to enter a value?

um we would likely have a different set

of code to do that. Um there is Python

code that can get a UI and uh we we will

see that um later on in the in the like

way later on towards the end of the

program. We'll see that um we can we can

write Python code to do a UI essentially

to to make like a almost like a web page

for someone to enter some input. We'll

see that uh much later on. So, we're not

going to get to that right now. It's

really complex.

Um what is the purpose of having

multiple cells? It's so that we can run

individual pieces of code within those

cells. It allows us to isolate, right?

Because I can run I can run code inside

of these cells and they don't affect any

other cell. So, it's it's just for like

debugging and isolation, which is nice,

right? I don't need to worry about

running all of it at once. I can run one

cell at a time.

Okay. Any other questions?

Um, can we execute multiple lines

together? Yes, we did that. Here I had

multiple. So, I'll I'll show you again.

I can do um print

um this is one statement

and then I can come down and do uh print

um this is another and then maybe I can

do some math.

So you can have as many lines as you

want

within a cell.

Within a cell, you can have as many

lines of code as you want.

Is there a way to tell it the order the

cells execute? Um, you no, if you if you

go up to um if you go up to run all,

it's going to run them all in order from

top to bottom. Uh, in order to tell

which cells to execute, you can

rearrange them. You can always like So,

I could rearrange these cells, by the

way, by I think there's a way to move it

down.

So, I can move it down. So now I'm

rearranging. So you can move cells. I

think you can even drag and drop them.

So notice how I took the one that's at

the very top and I'm moving it down.

Otherwise, you have to click, right? You

just have to like I can run them in any

order. If I just click like if I click

here, it will run that one first. If I

go back up here, it will run that one

next. So you just click around which

ones you want to run. Does that make

sense?

I can run them in any order as long as I

click on whatever order I want to do it

in.

Okay.

All right.

Perfect. So, that that wraps up that

demo. I hope it was informative. I hope

you saw the the print statement. Um

we're going to see that many times. The

input statement. Um that's pretty cool.

Um and you got to run you got to run

some Python. So, if it's your first time

ever doing programming, congratulations.

You ran some Python code. That is really

exciting. Um, so glad we got to do that.

Um, let's go back to our notes

and then we'll um

let let me share the screen.

Okay.

So the next thing on our agenda is to

cover variables and data types. So I

just said like text is the string data

type but let's learn about all the

different data types that are going to

be available to us inside of Python and

let's talk about variables. Um it's

going to be a good discussion. So um I

think what we'll do is we'll take a

fivem minute break now and we come back

and we can start this uh discussion

about variables and data types. Um, so

let's take uh a fivem minute break.

And so let's try to be back um around

uh 8:30.

Okay.

Try to be back around 8 8:30.

Okay. So, what are variables? These are

um

basically our way of storing data to

make it easier to reference them and

manipulate uh throughout our program.

So, we've actually already used a

variable. We we used one in our uh demo

we just did where we called the input

the name. Uh we stored that input into a

variable called name. And so um

variables just really are a reference to

some data. That's all they are. They

allow us to reference that data

throughout the program. We can store

information into a variable and then

access it throughout our code. Um so on

this screen are some examples of

variables. Now, variables have names,

which is why I said usually we want

those to be meaningful. Like X is a

valid name, but it's not that

interesting of a name. It doesn't give

us that information much information

about what it's what it really means.

So, probably not the best name. Um, but

we have things like uh we can we can

store some text inside of this variable

called name. We can store a number

inside of this um variable called price.

we can store uh a true or a false value

inside of this variable called

is_active.

Um and so variables will show up all

over our code and uh they are basically

our way to reference some values. Now

these things over here are basically

different types of data that we need to

learn about, right? So we need to learn

about what is a 10 versus what is in

something inside of quotes. is it a

string versus something that has

decimals which is a floatingoint number

versus something that is true or false

which is a boolean value. We need to

learn about those data types. But notice

how all of these are being referenced by

a um by a variable that has some name to

it. Okay. So the variable is this guy.

It is our reference to that data. Um and

we will use variables throughout um so

that we can have you know references to

information in our code.

So variables are fundamental um to to

working with Python.

Um now variables can store different

kinds of data. So I just alluded to

that. And so the different types of data

available to us in Python kind of fall

in these two different categories. one

being single values or what are known as

scalar values. So these are things like

integers. So the number 10, the number

1, the number 2,323,

those are all whole number integers. Um

floats, which are anything with a

decimal.

So 32.3,

3.14,

um 1.2, those are all floating point

numbers. Um, booleans only have two

options. They only have true or false.

So, they represent kind of a binary uh

value um which we say is true or false.

And um then we also have um complex

numbers which are which have imaginary

uh parts to them. We won't really be

dealing with complex numbers too much so

I wouldn't worry about them. But in

reality, Python supports working with

the uh complex numbers and doing complex

math. But uh so so complex numbers just

have kind of a real part and an

imaginary part to them. Um wouldn't

worry too much about that. Again, we're

not really going to work with those ever

throughout throughout the program, but

it does exist. Python supports it. So

scalar data, single values, think

numbers, think single numbers like

floats, think integers, um single uh

true or false values. So these kinds of

data can be stored into variables.

On the opposite end of the spectrum are

aggregated types that we are storing

multiple things.

So we're going to learn about all of

those, but um probably the most common

and one that we've already dealt with is

going to be a string. So a string is

technically an aggregated type because

it has multiple characters that form,

you know, an overall uh string, which is

a string is usually you you know it's a

string because it's inside of quotes,

right? It's inside of these double

quotes or single quotes. Um, Python

actually doesn't care about quotes

really in terms of if it's single or

double as long as you're consistent with

it. Like if you if you start with double

quotes, you should end with double

quotes. If you start with single, you

should end with single. Python doesn't

really care either way. Um, so strings

are going to represent um collections of

characters. Um we are going to talk

about sets which are basically like u an

array of unique values. Um so we'll talk

about sets we'll talk about lists which

are a really important structure. It's

basically an array that can hold many

different types of data. Um so we'll

talk about list. We'll talk about

tupils. So you if you see that word

tuple e that is um people some people

pronounce it tuple. I I like to call it

tupole, but um that is going to be very

similar to an array. It's just going to

have slight differences and uh if you

can change it or not. Tupils you

actually cannot change once you create

it. Um versus list you can modify list.

You can add things to it. You can remove

things from it. Tupils you cannot. So

we're going to learn about those

differences as we go along and start

working with these different types of

data.

Um but they are designed to hold

multiple values, right? So you can see

in that example that list has integers,

it has strings, it can it can hold

multiple types which is if you're coming

from other languages is generally not

the case. Um like arrays in Java, arrays

in C, they can only hold one type of

data in the array. They can't hold

multiple.

Um

so uh then finally a dictionary. A

dictionary is if you're coming from

other languages, it's like a map, a

hashmap. Basically it allows you to have

uh keys mapped to values. So it's a

really dictionaries are highly useful

for storing information where we want to

reference like this value maps to this

value. So for instance in this

dictionary the string a maps to one and

then the string b maps to uh you know

two and or whatever it maps to. And this

will allow us to look up values in the

dictionary. So we could look up, hey,

what is the value stored at key A or

what is the value stored at key B? Those

kind of things. Dictionaries will be

incredibly useful. We're going to

explore all of those more as we go along

in the lesson, but um for right now, it

should be making sense that there are

some data types that store multiple

values like array or sorry, lists, um

dictionary, strings, and then there are

some data types that only have a single

value like a single number like a float,

integer, um boolean.

Okay, so more to come on aggregated

data. We're going to work with those,

learn about the differences, learn about

what it looks like in the code to work

with the set, a dictionary, tupil, list,

but those generally hold multiple values

or can hold multiple values whereas um

scalar data is only going to hold one.

Okay.

All right. So uh so as we said earlier

um you know integers, floats, booleans,

they only hold a single value. By the

way, inside of Python, if you ever want

to see what the type of a variable is.

So let's say we know we have a variable

called name. We can always check the

what data type it is by by using the

built-in type function. So we can use

type and then pass in that variable

and this will display what data type it

is. So um if we stored the value 42 in

some variable called int, if we um

displayed if we did type um if we did

type of this it would uh produce int

which would say okay this value is an

integer versus 3.14 that's going to be a

float versus capital t true that's going

to be uh the boolean type bool.

Okay. So, uh we have integers, we have

floats, we have booleans, all of which

we will use throughout and we'll see

where we will use those one versus the

other. We'll learn about that.

Um as I said, complex. So, uh just

showing you here that those exist

obviously. Um like I said complex has uh

a real part and an imaginary part which

you can access separately. So if you

store a value as a complex you can uh

access its real and imaginary parts

separately which you may need to do for

some type of uh calculations.

Um again we won't really work with

complex numbers in in this program. So

not a big deal for us but it is

supported

and you know a lot of um mathematical

packages in Python will use complex

numbers uh if they need to but we won't

really do it in this program. There's

not really a need to for us.

All right. So aggregated data um we have

those strings which we've already seen.

Those are the things inside of quotes.

We have sets which are going to be

collections of data um that are unique

basically only allowing one uh copy of

those elements inside the set. We're

going to learn about that. Um list which

is going to be a collection of items

which we can change, we can add things

to it, we can remove. Um lists are

really awesome uh structure in Python.

Um, what I want you to see right now

though is you can start to see the

syntax differences, right? So, like a

set, um, a a set is where we have, uh,

this brace. Notice that a set is created

with a curly brace versus a list which

is created with a bracket. So, right

away, like when you see a brace, you

should be thinking either a set or

dictionary. Those are the two things

that are created with a curly brace. Um,

and you know it's a dictionary because a

dictionary will have the colon which

will map I'll show you that on the next

screen. But that will map things from

key to value. Um, depending on if you

know left and right of the colon. Um,

but do you guys see that like the syntax

difference of a list? A list has a

bracket set has a curly brace. Um,

that's just one small difference. you

know, we're going to learn like what is

the actual difference between a set and

a list, but that's just one I'm pointing

out right now.

Um, what does mutable mean? So, mutable

uh means that we can change it. It's

it's able to be changed. So, immutable

would be we cannot change it.

Yeah. And and one thing about a list

that's really nice is every list has a

natural ordering to it which is actually

really beneficial. So a list has a

notion of the first item, the second

item, the third item, the fourth. That's

really important for accessing data

within the list. Okay. So lists are

really powerful. Um

yeah. So a a set the reason it shows

it's in a different order is because a

set does not maintain order. A set never

maintains order because um it a set is

not you do not access items by by order.

So that's just something unique to a set

is that it doesn't have a natural order.

So every time you print it out, it will

display in a different order.

Potentially it's random. It's random

order when you when you display it. A

set is just meant to be a general

collection. Think of it like a bucket.

Like here's this bucket of items that I

have.

It's just a collection of items. A list

actually maintains an order, a

consistent order of items.

This is different.

So we'll talk more about that when we

get into those.

Okay.

All right. So I wanted to show you also

the tupole in the dictionary. So a

tupole

is also a collection of items. Now the

tupil is ordered. So it's like a list.

It's ordered but it is immutable.

Meaning you cannot change a tupole. So

once you create a tupil you cannot

change it or else you'll get an error.

Python will tell you hey this is

immutable I can't change this. So if you

try changing being if you try to add

something to the tupil if you try to

modify one of the entries in the tupil

like if I try to if I go in and try to

change this a um to a d

um this would not be allowed. This would

this would throw an error. The

interpreter would say hey you're trying

to change something that cannot be

changed. So tupils are immutable but

they have a benefit beyond a set of

actually being ordered. So there there's

a natural ordering to a tupil where this

is the first item, this is the second

item, this is the third and every time

you display a tupil will be in a

consistent order. But tupils are not

like a list. You can't change it. So

tupils are useful for situations where

you want ordering, but you don't want

anybody to change any of that data

that's in the tupil. It's it's not

changeable.

Mutable meaning it just means changeable

like you can modify it. If if something

is mutable, you can modify it.

Immutable like a tupole is immutable. We

cannot modify it once we create it.

That's what it is.

Okay.

Yeah. Okay. So then finally a

dictionary. Now you by the way um look

at the tupole. See how it's created with

a parenthesis.

So that's different than the curly

brace. That's different than the

bracket. Right? So a tupole you know

it's a tupil because of the parenthesis

and the items are separated by a comma

just how just how they are in a set and

just how they are in a list. Um

so so the the parenthesis gives it away

that it's a tupole. Um now look at the

dictionary and the dictionary is um a

collection of key value pairs. So this

is a key value pair. This is a key value

pair. Um this is a key value pair and on

and on. We can have as many as we want.

And one thing I want you to notice about

this is there is no restriction on the

data types of the keys and the values.

So keys can be integers, keys can be

strings, values can be integers, values

can be strings, values could be floats,

values could even be other dictionaries

or lists. So value like we could have

what's called a nested dictionary where

we actually have something mapping over

to another dictionary.

That's totally possible in Python. So we

can have dictionaries that part of the

values inside of the dictionary actually

have our dictionaries themselves and

that would represent kind of a nested

structure there. So for instance this

name could map to a dictionary with

everybody's name in it. Um or it could

map to a list um you know

could map to a list it can map to

whatever it could map to a tupole. Uh so

you there's really no restriction in

what the keys and values uh are going to

be.

Uh Brent is it more efficient than the

other uh is what more efficient than the

other methods? Just want to clarify your

question so I so I answer it properly.

The tupole versus using an array.

Yeah. Yeah. Yeah. So uh these are all

good questions. So um the tupole

is guaranteed not to be changed. So it

is a little faster when we are looking

up items like when we are referencing

items. It's a little bit faster because

uh we know that it's not going to be

modified ever. So everything is going to

be consistently in the same spot. So

like whatever's first is going to stay

first, whatever's second is going to

stay second and on and on. So tupil is

is nice in that sense. A list can be

changed. So whatever is first may not

guarantee to be first in the future. We

can modify it. We can remove things. We

can add things to the list. So we can

expand. The list is very like dynamic.

The list. So the list is less efficient

because it's way more dynamic. Does that

make sense? Like it can change. You can

add you can keep expanding the list by

adding things to it. You can shrink the

list by removing things from it.

So list is way more dynamic which for a

lot of scenarios is useful,

right? We want to be able to add and

remove and modify things.

Um but a tupole is more rigid in the

sense that once you create it, you

cannot change anything about it.

Yeah. Yeah. So a dictionary is good for

Yeah. Like a phone book would be a good

example of a dictionary because you with

a dictionary you're usually looking up

things. So you so like a diction in a

phone book you have a name that maps to

a phone number.

Um so yes you have a you have that a

dictionary will map a key to a value

just like a name would be mapped to a

phone number. So yeah a phone book makes

a lot of sense.

Um,

a list, a list is like any is like a

normal like like your grocery list. Like

you may add things to it, you may remove

things from it, you may change things on

it. It's very dynamic. Um, a tupole is

kind of like a fixed um set of data

that's ordered in some way. So maybe

like um what you would see on on a on a

letter like you have your name, you have

your address, you have um your zip code,

like you kind of have those and it it

should stay that way in order to mail

the letter kind of thing.

Uh can you convert a tupil to a list?

Yes, you can do vice versa. You can

convert a tupil to a list and you can

convert um you can convert a list to a

tupole. Yes, you can convert between

them.

I'll show us examples of that later.

Okay. So, just to recap there,

tupil is not changeable, but it has an

order. So, it has a natural ordering to

it. Whatever is first is first. Whatever

second is second, third and third. So,

you can access things based on their

position within the tupole. That's

really nice. But you cannot modify

anything about a tupil once you create

it.

Okay. A list has an ordering to it. You

can access things based on their

position. But a list is dynamic. It is

mutable. Meaning you can change it. You

can change values. You can add things to

it. You can remove things from it. Okay?

So very dynamic. That's what a list is.

Um dictionary. It maps keys to values.

No restrictions on what those keys and

values can be.

All right. And then a set. A set is

think of it like a bucket. It just has

things in it. A set has no order to it.

So you cannot access things based on

their order. And every time you uh

display the set, you can get a different

ordering. Um

but a a set only is special in that it

only allows unique items. So if you try

to put multiple copies of a piece of

data, it's only going to keep one of

them. So a a set is like a bucket with

only unique things in it.

Okay. So and sometimes that's really

useful is to know like what are the

unique values? Uh a set would help us

maintain that.

Any questions about those? You know, we

have to we have to work with this and

see this in the code and we will. But

just any questions right now about these

different types of data that we're

talking about.

Okay.

Very good.

All right. Let's talk about assignment.

So, what that means is um Oops. Let's

talk about assignment which means that

we will be um taking a variable name and

assigning data to it. Now we've already

seen this. We already saw it in our demo

where we did input. We did name equals

input.

So the equals symbol is how we assign

values to a variable.

That makes sense, right? It's very like

self-explanatory.

But um what we should think about with a

variable is really the fact that a

variable is is a reference to that data.

Okay. So when we say x= 34, we are

assigning 34 to the name x. So x becomes

a variable which is referencing the data

which is an integer 34. Right?

What's really interesting about that and

this is how you can kind of test your

intuition of the fact that this is a

reference is if we come along and have

another name Y and we set that equal to

X.

This is just saying that we are creating

another reference that is equal to the

reference we already have. Now, why

would we ever do that? Probably we

wouldn't. That's kind of redundant. But

this just proves that they're ultimately

references because when we display x, we

get 34. Of course, that's what we stored

the value 34

uh referenced by x.

And then when we print y, we get the

same number, right? We get 34. And why

does that happen? Because we we

literally declared y equal to x. Meaning

y should reference the same data that x

does. Okay, so as variables they are

equal meaning that um X is being

assigned to Y meaning Y should reference

the same data that X does. So they they

uh contain the same data. Now what's

interesting is if you print out the ID.

So the ID is the internal

um the internal memory address

of of the reference.

Um now it usually we don't care about

that but this is just to prove the point

is that you can see these are the same

address. These are the same. That's by

design because we're saying okay I have

this reference X which is referencing

this data 34. it's stored at this

address. Um, and then when I come along

and say, okay, y equals x, that's just

the same reference. You see how it's the

same exact address,

same reference.

So, just proving that variables are

literally just references to data. They

allow us to reference that data, which

is really, really, you know, nice. So,

we can reuse x throughout the code. Um,

we can reuse name. we can you know

whatever we create we can reuse.

Um if you look over to the right we have

an alternative example which um now

resets y to store a new value. So

instead of saying y equals to x we

actually overwrite y and reassign it to

the integer 78. That's a new piece of

data right 78. So now if you look at

their their uh references, they're

different. These are different. And that

makes sense because now they're pointing

to two different uh pieces of data,

right? X is pointing to 34. Y is

referencing to 78. So of course they're

going to be different uh different

addresses. And this is a bit of a typo.

This should say ID of Y

because we're ref we're talking about Y.

It's a bit of a typo there.

Okay, so hopefully this now this example

is just to reinforce the fact that when

we use the equal sign, we're setting

equal we're setting a variable name

equal to a piece of data, right? And

that is creating a reference to that

piece of data.

That's all we're that's all we're saying

with this. So we are assigning a piece

of data to that reference X or Y or

whatever it is.

Okay.

All right. Let me ask you guys. Um, what

is the default data type of a variable

assigned using the input function? This

is an interesting question. We didn't

actually cover this, so I'm really

curious to see what you guys think about

this.

A lot of votes for for string.

Let's get a few more.

Perfect. Yeah. So, water votes receipt

it is a string. So, that that begs the

question like what happens if we input a

number? Like what happens if we put in a

two? What happens to that? You know that

two will actually be read in as the

string two. So it would be So if we use

the input and we it pulls up that text

box and we put in a number like two

um and we set that equal to the variable

x, whatever we name that name x,

whatever. What that really means is x is

going to be um equal to the the um x is

going to be equal to the

uh string 2. So that's something to be

cautious about with the input is it

always assumes the input data is going

to be a string. So luckily there's a way

to convert between strings and numbers.

So if we wanted to turn this into the

actual number, what we would do is use

the the data type function int, which

would convert uh this would convert it

over to the numerical two. Would

actually convert it from a string to an

integer. We just use int. Or we could

use like if we if somebody put in a

decimal like 2.5

then um we could do a float

of 2.5

and that would convert that over to uh

the the number.

Okay, let me actually show you guys

this. Let me go over to Collab real

quick and show you guys this. I know

it's not in a demo, but I think it'll be

better if I just show you what I mean by

this because this is an important point

with input.

So, let me uh stop sharing there. Let me

go over to Collab for a second so I can

show you literally what this means.

So, go back into the notebook here. So

what I want to show you is that um when

we do input

the default type

is string.

So for instance when I do um

when I do uh uh value equals to input

and let's try um enter your age.

Oops. Enter your age.

And then we uh run this.

So we enter the age. Now this is going

to be read in as a string. So even

though I'm putting a number there, it's

actually going to be read in as a

string. So now

look at what the type of value is.

It's a string. Do we see that? So

string. So this number even though we

put in a number it gets it the the input

function always converts it to a string

no matter what we put there. If we put a

decimal if we put a a large number it's

always going to assume it's a it's a

string. So luckily

um we can convert to an integer

by using int the int function.

So um we can print sorry we can say

value

uh or we can do int value which which

will convert that 32 string because

right now if I were to um just display

value it's a string 32. You can see it

inside of the quotes. But now when I do

this uh and I can run that now it's an

integer. Do we see that now it's

actually a number

which is great. It no longer has those

quotes. It's actually going to be

treated as an actual integer which which

may be useful for calculations or

storing it or whatever whatever we need

to do with it. So that's just one piece

of caution with the input is if you're

working with numerical data it's going

to treat it as a string. We have to

convert it.

Okay

questions on that. Does that make sense

to us? like the input's always going to

accept the input as a string. So if we

want to work with it alternatively

um we should convert it.

Uh you can yeah so like you could

convert um if I did this if I wrapped

this around in the int function that

would automatically

take whatever we put whatever this

returns would automatically be um cast

over to an int. So we could do that. So

let me show you that. So when I run

this, I can put in 32

and it it's like automatically going to

be casted to an integer. So there now

it's an integer. Does that make sense?

Like when I wrap this int around the

input, it's going to automatically

convert

Uh, what did you put in the input box?

So, yes, you'll get an error if you

don't put in a valid integer.

So, let's put in like if I put in my

name,

this is going to be this should be an

error because I don't know how to

convert this string over to a number. It

doesn't make sense to do that, right?

So, this should be an error.

Right? That will be an error because

it's a string.

So why did you get an error? Uh input

enter your age value int value.

Uh did the did the text box show up?

Maybe try separating it into a different

cell.

Try try putting the other two lines in a

different cell. Um, you need the text

box to show up and then you need to

enter something.

Yeah.

Okay.

All right. Does this all make sense? Any

questions about this? About the input

function.

Okay.

Good. Okay, let me go over back to the

notes then.

Okay.

All right. So, we have another demo.

We'll do that now. Uh I was just kind of

doing one, but let's go back over to

this will be demo five. Let's do that.

So, we're going to practice assigning

different values um to variables and

displaying them just so you get in the

habit of being able to create your own

variables and just go through that kind

of one more time. We'll do this one

relatively quickly um and then uh move

on.

So, this will be uh demo five.

So, let me pull that one up for you

guys.

Okay, let me share my screen.

All right. So, this is going to be demo

five. Um,

now again, like feel free to use

whatever platform you've been using.

Collab, Jupyter Notebook. I know this

instruction says set up a Jupyter

notebook. Feel free to use whatever you

want. You can use Collab. Um, whatever's

been working for you to build your build

your notebooks. So obviously this this

looks a little different than collab but

it's because it's the Jupiter. Um so we

create a notebook.

Now what I want you to see

is this takes the approach of everything

we just did. Let me zoom in on this. Uh,

I know that's a little small,

but this is doing everything we just

said we could do where we

um essentially take

So, I just want to zoom in on this. Um,

notice that we

uh take the um input and this will be

saved as a string.

Um,

so this will be saved as a string and

this will be saved into this name. And

for instance, this will be saved as a

string, but we convert it over to an

int, which is exactly the kind of

example I just did, right? Where we take

take an input, we convert it over to to

uh int.

Does somebody have Yeah. Does somebody

have the demos available? like if if

somebody doesn't mind sharing those in

the chat. I again I don't have the PDFs.

They should be from your LMS. They

should be in the reference material.

There should be a demos folder that you

can download. If somebody has those and

doesn't mind sharing them.

They have that folder of them, like a

zip folder of them, that'd be fantastic.

Yeah. Thanks. Thanks. This is This is

the demo we're going through currently.

Perfect. So, for you guys having trouble

navigating the demos, please download

this zip folder.

Download the zip folder that that these

guys are uploading. Thank you so much.

Download the zip folder so you have all

of them.

Please take a moment to do that.

Okay.

Uh,

copy the code and got an error at height

value. Use foot, not meter. I mean, it

shouldn't matter. It, you know, you

should just be the point of that one is

to put in a decimal.

How to create a new file. Um, what

platform are you on? Collab.

I don't know what platform you're on.

Collab. Uh, just go to file, new

notebook.

New notebook in drive, I think is what

it's called.

Do you see that? It should be like it

should be at the top. There should be a

file and then new notebook.

Let me go over to it.

Uh,

this one. You don't see this

file. It's at the top. The top of the

notebook. Do file and then new notebook.

You don't see new notebook.

Uh if you if you don't see that, just go

to a new tab. Just go to a new tab and

go to um Google Collab.

You can always do that. Just go to just

start a new um just go to Google Collab

and then it will let you like launch a

new notebook. So just just do that. Just

do a new tab if it doesn't work.

Okay.

So, by the way, one of those examples

was entering a float. So, it looked kind

of like this. So we had um our our

height is equal to float and then we had

uh input and then we had um enter your

height and then this was um uh some sort

of uh this should be some sort of

decimal value. So let's say it is um I

don't know uh 5.7

whatever that is uh feet it doesn't it's

just some decimal um and then we hit uh

we hit enter that will store the height

as a float so that when we um display

the height uh it will be rendered as a

float appropriately right that's what

that that's what should happen

that's the point of that It just needs

to be some decimal. It should work.

All right, let me go back to the demo

document.

All right, were you guys able to run

some of these? Like, were you able to

run some of the inputs and change them?

So, try these out on your own real

quick. like try doing int and then input

for enter your age. It should convert

that. You should be putting in a number

or else you'll get an error and it

should convert that over.

I by the way I wouldn't worry about this

last one uh because we haven't learned

about the comparison operator yet which

is this equals equals. So we'll learn

about that in a in a little bit in a few

minutes. So, don't worry about that one

too much right now. But at least these

first few should make some sense and we

should be able to do.

Were you guys able to run one of those

and convert over the the float or int

and do the input and convert it?

Did that work for you?

Give it a try.

Let me clear that. Any questions about

that?

Should look something like this.

Good. We're good on that on converting

over the input. Okay, perfect. Sounds

like Sounds like we're able to run that

and uh it was okay.

What are you entering for the feet?

Like, are you literally entering like

quotes?

Yeah, that's not going to work when you

do that because it's going to um there's

a string f, there's a character there.

it's not going to be able to convert

over to.

So if you did if you did 6.4 that would

work.

Any decimal should work. But like the f

is a character. So the the float doesn't

know how to convert over a character,

right? Yeah. So so that's not going to

work. You need to put in a decimal to to

be able to convert over to the number.

Okay.

Very good. Very good. Let's go back over

to our notes so we can continue along.

1.7. Yeah. If you have any if you have

any character, it's not going to work.

It's not going to work. You need to put

in you need to put in a decimal.

All right, let's talk about operators.

So, these are going to be really

important. Um,

let's talk about operators so that we

can uh

uh be able to compare things and work

with things. Um so let's let's talk

about Python operators.

So what are operators? What do we mean

by that? In Python, operators are

special symbols or keywords that perform

operations. So as the name suggests,

it's performing some level of operation.

Um which means that the interpreter

should do some sort of logical

operation, mathematical operation,

relational operation to produce a

result. Um, so usually that means

there's going to be multiple variables

that are going to be used to do some

operation between. So an example of an

operation would be like adding,

subtracting, multiplying. That's an

operation. But we can have logical

operations like taking the um logical

and or logical or of things. We'll see

what that means. But um in Python,

there's many situations where we want to

we want to be able to do operations

between variables. Whether that's simple

mathematical or maybe some type of

relational like testing if a value is in

a list. That's an important operation.

Is 10 in my list? Is five in my list? Um

those are important operations. So we

want to learn about these operators and

they're going to be really important for

us going forward is because these will

be very standard. um things we will use

as we uh go along. So we're going to

spend some time talking about operators.

Um so it turns out in Python um you can

kind of group operators into many

different categories. Um there's going

to be standard arithmetic operators.

Those are your everyday things like

plus, minus, um division,

multiplication. Um, assignment

operators, which we've already seen, is

things like equals, where we're setting

a reference equal to something. We've

already seen that. That's an assignment.

Comparison, which is things like greater

than or less than. Those are important

for comparing values, comparing

variables. Um, logical operators are

going to be something like and and or,

which will um do a logical operation

between two two boolean values. That'll

be important. And then we have a

collection of miscellaneous operators.

Um those will be things like is

something in a collection like is five

in a list? That's an operator. So we'll

talk we're going to talk about all of

these but just pointing out that there's

many different categories of operators

in Python.

Okay, let's first talk about the

arithmetic operators. So these are going

to be your standard everyday um uh

operations between numbers. So if we

have numerical values like integers or

floats, we can do math between them.

That makes sense. Like that should be a

capability of Python and it certainly

is. We can add things, we can subtract

things, we can multiply things, we can

divide things. So um here are all those

operators. We have plus minus the

asterisk is a multiplication. So x

asterisk y will multiply those together.

So if we have two variables, one of them

is 50, one of them is four, we do x

asterisk y, that's going to multiply

them together to get 200. Pretty pretty

straightforward. Um

division is one that we should be

careful of. Of course, like we don't

want to divide by zero. So if you I if

the uh this secondary value that we end

up dividing by is zero, that'll give us

an error. Um the interpreter will say,

"Hey, you're trying to divide by zero."

We can't do that. It'll it'll produce an

error. So that's the only thing we have

to be on the lookout for with division.

Just don't want to divide by zero.

Um

so all these are pretty standard. I

think they all make sense.

Hopefully they do to you. I think

they're all pretty standard. you know,

the kinds of things you'd see on a on a

basic calculator. They all make sense.

They should exist. Now, here's some more

exotic ones. Um, I don't know if you

guys have ever seen the the modulus

operator, also known as modulo. This is

one that returns the remainder of a

division. Okay? So the the percentage

sign is a mathematical operation between

two numbers that returns not the

quotient like not the actual division

result but the remainder. So 50 divided

by four

um you know four goes into 50 um it goes

in there uh uh 12 times evenly but it

has two left over right. So there the

remainder there is two. So the result of

x mod we would read this as x mod y or

modulo y um returns two. So if you're if

you're unfamiliar with the modulo

operation that seems a little bizarre

that you take these two numbers

um oops it seems a little bizarre that

you take these two numbers and you like

do this operation and you get a

remainder result but it's actually a

very powerful operation. Um the reason

being is that sometimes we want to know

what the remainder is more than we want

to know what the quotient is. For

instance, things that are very like

cyclic in nature. Um so maybe we cycle

through a collection and we want to know

like how many times do we cycle through

and then we have something left over

which is the remainder. Um so the modulo

operation is pretty useful. You could

also check like if a number is even or

odd using this. Like so if you modulo by

two and it returns zero, that means it's

even, right? Because that means there's

there's nothing left over when I divide

by two. So modulo is kind of a nice way

to check if a number is even or odd. Um

so modulo is a pretty nice uh operation.

We'll use it from time to time. Uh but

that is the percent operator. So x

percent y will look for that remainder

of the division. Um now there is also a

double slash operator which is the

integer division operator. This is kind

of the reverse of modulo. It takes the

largest integer quotient that that uh we

can do from a division perspective. So

remember I said 50 / 4. We can divide 4

into 50 12 times evenly and we have two

left over. So the integer division will

just return to us an integer always

which will be that quotient.

So this is the quotient

um and this is the uh remainder of 50 /

4. So the integer division returns to

you that whole number like the largest

number of times that that number goes

into the other. So 12 times evenly

obviously there's a remainder there but

um but but yeah so integer division that

one's useful if we want to know like how

many times can I fit a value into

another value a whole number of times

and that happens from from time to time

we may need to know that.

Okay last operation here is exponent. So

the exponent is the asterisk asterisk

operator. Um so that raises a number to

a power. Um so for instance like x star

y or asteris y would mean that we are

doing an operation like 5 to the 4th

power um which is 625.

Okay. So asterisk pretty useful. Like

probably the most common asterisk would

be squaring something which would be x

um star star 2 which would would would

be um x squared. So I mean that's a

pretty common operation there is to

raise something to the second power

maybe the third raising something to the

fourth probably less common but um the

the asterisk asterisk operator is is how

we do exponents in Python.

Okay,

so these are all basic arithmetic

operations we can do between variables

in Python. All right, any questions on

those? Do those kind of make sense to us

from a syntax perspective?

Pretty straightforward, I think.

Hopefully nothing too surprising there.

Um,

do you used to use module all the time

for date date calculation? Yeah. Yeah.

like when you uh find out how many like

days how many weeks or where you are in

the week, you cycle through like uh

modulo 7 or something.

That make sense?

Okay,

very good.

Okay, I want to talk about assignment

operators now. Now we've already seen

this which is the basic equal sign that

is a data assignment operator right so

that means that we are setting a value

equal to a reference so we are storing

data inside of this reference variable a

we use the basic equal sign as our

assignment operator so that equal sign

is called the assignment operator now

what's really awesome is we can combine

this basic assignment operator with our

arithmetic ones to update values

um and modify them uh as kind of a

shortcut to say uh so so for example

like a plus= 5 really represents the

fact that I want to reassign a to the

result of a + 5. So this means take

whatever it is add five to it and

reassign it to the value of a. So this

is the same thing as if we just shortcut

it in and Python will recognize if we do

plus equals 5 it's the same thing. So

and actually we can do that with any of

these arithmetic operators. So if we

want to take a variable multiply it by

two and reassign it to that variable we

can use star equals. So like a asterisk

equals 2 is the same thing as if we were

to reassign a to the value of a * 2.

Does that make sense on the

reassignment portion of that? So plus

equals divide equal modulo equals star

star equals would exponent something and

reset it back to the variable.

um minus equals we'll subtract and

reassign that back to the variable. So

you know x minus equ= 3 we'll subtract

three from x and re and basically update

it right reassign it back to x.

So so that's pretty useful like whenever

we need to do an operation and add like

um you know a a very typical

reassignment is to do like a plus equals

1

That's a very typical reassignment

because what this is the same as is a

equals a + one. So that's like a single

increment of a. We're just updating it

by one.

So plus equals 1. We may see that from

time to time.

A loop coming on. Yeah. Yeah. These are

used in like while loops. Yeah. Like you

do plus equals and you increment it

until you reach a certain condition.

Yeah.

Now, if you're coming from other

languages, if you have programming

programming experience, you're coming

from other languages, Python does not

have an increment operator like plus+. I

wish it did, but it doesn't. So, like I

know in in Java and I think C they have

um you can do like uh a plus+ or

actually reverse you can do plus a but

um that does not exist in Python

unfortunately. You have to do the plus

equals reassignment. So they don't have

an increment operator. You'd have to

you'd have to do just plus equals one to

do the same effect as plus+.

So I I know some people ask about that,

but yeah, doesn't exist unfortunately.

All right. Any questions about

assignment? It's just really the equal

sign and we can tack on the arithmetic

to do some type of basic math and

reassign to the variable.

Hopefully the fact that we're using a

single equals makes sense. Where people

get confused all the time is the

difference between a single equal sign

and multi and two equal signs which

we're going to see. Two equal signs

means something completely different

than a single equal sign. Single equal

sign is an assignment. We are taking

data and storing it in a reference

variable,

right?

But multiple equal signs, we're going to

learn about what that means. That's

actually a comparison.

It's something different.

All right, we'll continue. Thank you

guys. All right, so we're talking about

uh comparison. So, uh we're going to

talk about a few operators that allow us

to compare two values. Now, this is

going to be useful as we go forward

because sometimes we want to know when

is a value bigger than something or less

than something or equal to something,

not equal to something. Those

comparisons are going to be useful.

um and we have a collection of operators

to do that for us. So again, one that I

think a lot of people get confused on is

the um equals comparison operator which

is uh the double equals symbol. So a lot

of people get confused on that. What is

the difference between a single equal

sign and a double? This double equal

sign is checking if two values are

equal.

Um

so for instance we have uh these two

numbers x and y they're both integers

that are 20. We check if x equals equals

to y and that returns true because

uh these two values are the same. They

both equal 20. So when x equals equals y

that is a true statement. So these

that's something to realize is that

these comparisons are things that return

booleans true or false because a number

is going to be bigger than another yes

you know true or false they they are a a

uh comparison that gives us a kind of a

yes or no answer. Um

so the equals equals checks if two

values are the same and then the uh not

equals operator which is uh an

exclamation point with an equals um

checks to see if two values are

different. So they are not equal. So for

instance if we had um uh 45 and 24 we we

uh do x not equals y that would return

true.

Um now if we had these two values as

before and we checked here x not equals

to y um this would be false because they

are equal right so um

not equals to checks if values are

different so that's a simple comparison

are they not equal um so so in this case

that would return true

so these are pretty useful if we want to

compare directly is a value equal to

another we use the equals equals If

they're different, we use the not

equals. And we're going to have

different scenarios where we will use

those.

I also want to call out the basic, you

know, greater than and less than. So the

this first one is the less than

operator. It is uh going to be obviously

returning true when a number is less

than another number. So when we have

things like 20 and uh 30, this x less

than y would return true because 20 is

definitely smaller than 30. So this

returns true. Um

and then greater than checks if a number

is bigger than another. So that

comparison uh x bigger than y in this

case would uh return true as a

comparison. So again these are all

operators that check uh comparison

between two numbers that will be uh

really useful as we go forward and start

to work with data and numbers and we do

comparisons.

uh we will do those all the time later

on.

Now there's also scenarios when when we

want to know is it less than or equal

to. So that operator just tacks on an

equal sign. So less than equals

is the less than or equal to operator.

So for instance 10 less than or equal to

30 that is true um because 10 is

certainly smaller than 30. But um it

would have been true even if x was 30.

That would also be true because 30 uh 30

equals to um 30 would equal to 30. That

would be a true statement.

Um greater than or equal to same same

scenario. We have a greater than and

then we have an equal sign right after

it. This returns true if something is

bigger than or equal to another number.

So here's an interesting one. We do 30

bigger than or equal to 30. That returns

true because 30 equals to 30. That that

makes sense. So less than or equal to

bigger than or equal to we can we can do

with these simple operators.

Uh is greater than greater than similar

to the usage of brackets? No.

Uh so greater than or greater than is

what's called a uh a bit shift operator.

It's a little bit different. Um I I

would I'm going to save any explanation

that just just look that one up is what

I'll say. It it does like a bit um a bit

manipulation which is um a bit of a bit

of a hassle to deal with but we we won't

ever use greater than or we won't ever

use greater than greater than. It's it

does some sort of a shifting operation

like a bit mathematics which we we don't

need to do.

Okay. So those are comparisons. Um let's

look at our logical operators. So now

these ones are going to be really really

interesting and useful when we get into

controlling the flow of our program. Um

so logical operators are used for

combining conditional statements. So

conditional statements are things that

return these are statements that return

um true or false. So they return a

boolean and we can it's it's like we are

combining them together in certain ways.

Okay,

so the and operator, let's look at that

one first, which in Python is the

literal word and. So that's very nice.

It's it's literally the the keyword and.

Um, and what this does is it takes the

result of some boolean comparison and

some other boolean comparison and

returns true if both of them are true.

So and will only return true as a

combination if both individual

statements are true. They both have to

be true. The moment one of them is false

and will return false.

So this is useful for doing a

combination of things where we want

every individual thing to be true. So a=

1. This is a true statement because a

equals 1 and then b= 2 is a true

statement. So both of these would be

true. So therefore when we combine them

with the and this overall combination is

true.

So keep that in mind. These operators

are ones that combine individual logical

statements or conditional statements.

Right?

Okay. Now the one that is less

restrictive than and is the or statement

which um is used when you only want at

least one of the statements to be true.

So if we want to combine these things

and only require at least a minimum of

one to be true, we use the or statement.

So for instance, a= 1 is true because a=

1. So that's true. and then B equals

equals to 2 is false. So this one is

false. But that doesn't matter from the

perspective of or because we have a

minimum of one of these statements being

true. So or when we use the logical

combination of or um we just need either

or to be true. So a= 1 is true. So this

overall returns true.

So or is something that will combine

conditional statements and return uh

return true if at least one of them is

true. If all of them are false or it

would return false because none of them

are none of them would be true.

Okay. So we have and we have or and then

we have not. So not is an interesting

one. Not essentially reverses a boolean.

So if we have a statement that is

inherently true and we put a not in

front of it, it will invert that to be

false. If we if we have something that

is false and we put a not in front of

it, it will return uh true.

One of the interesting examples in

Python and this trips up people all the

time is the fact that Python treats zero

very specially. So the integer zero

is

oops the integer zero is inherently

treated by Python as false.

So Python treats zero as false and then

every other integer as true. Basically

being Python is indicating that it is

something that is not zero. uh anything.

So like um B equals to one would be

treated as true because it's as long as

it's something that's not zero

then Python treats that integer as as a

true boolean essentially. Um so why

that's interesting is if you put a not

in front of this this would actually

return true because not false what is

the opposite of false? It is true right?

So not false it would return true. So we

we will see not from time to time. Uh

not shows up when we want to negate

something. So when we um you know you

know maybe we have a an iteration an

iterative loop and we say while not

finished and we you know then we will

execute a bunch of statements while we

continue to not be finished and then the

moment that that it finishes then it

then the loop would be over. So not is

powerful to kind of invert uh trus to

falses and falses to true.

Um so maybe we want to check if

something is not empty. Meaning that um

if it's empty

uh if it's not empty that would be

false. Not empty um you know maybe it

would return true. So not is something

we will uh see from time to time as a

negation operator logical negation.

Okay.

Any questions on

uh any questions on these operators?

These now these we're going to use these

in the control of the flow of our

program.

One more example for not. Yeah. So a

pretty typical case for not would be

something like this where um

uh maybe we have some code oops I always

forget to swap over to this maybe we

have some code that checks so if we have

a list

if we have a list and it's uh empty

okay and it has nothing in it let's say

it has nothing in it we could we could

have some code that says like if um if

not

uh list

um then we then do something. So then if

not list uh meaning that it's not empty

then check then grab the first value.

Let's say that grab the first value. So

we'd have some code like this. So uh

this this not is used to like we can

negate the fact that this is going to be

empty and then this would be true and

then we can continue to access something

because that would mean it's not empty.

So not empty is a pretty standard use

case for not like to check that

something is not empty.

Uh can we also use not for checking

value in the list? Yeah, that's that's

what we're doing here to say like is it

not empty?

Okay.

All right.

Let's go to some miscellaneous

operators. So we have now some of these

are going to be incredibly useful. The

one on this page not that useful. The is

mainly because it's very rare that we

would uh that we would check these. So

is is what we call the identity operator

and this is something that um checks to

see if two references are the same.

Okay, two references are the same. um

meaning that they're referencing the

same piece of data. Um so now this is a

very interesting case where we have a

equals to a list b equals to the list.

However, when we ask the question a is b

this would actually return false. Now

that seems very counterintuitive but the

reason that's the case is because we are

creating two different references. We're

saying A equals to this list, B equals

to this list, which is a whole new piece

of data.

It's a whole new piece of data. So

therefore, we can't claim that they're

the same reference even though they're

the Now what would be true is A equals

equals B because their data is the same.

That would be true, but their their

references are different because they're

different variables, right? A and B are

different variables.

Different memory location. Exactly.

Different references. So A and A is B is

the same as checking um if ID A equals

equals IDB. Does that make sense? That's

basically checking that that logical uh

comparison if their addresses are the

same. It's the same check. So is is

basically a shorthand for doing this.

And so th those would be false because

they're going to be two different

references, A and B.

Um, however, we can use our not. So A is

not B. That's actually true because it's

the inverse of is, right? So that that

actually would invert the false and this

would be true. A is not B. That is true.

Behind the scenes, yeah, like the

memory, yeah, the location in the

computer memory is different. Yes,

because they are different variables,

different references.

Yeah, the data is equal. The data is

equal, but the references are different,

which is what this checks. You know, we

have two different names, A and B. Those

are different.

Now, take a look at this last example.

This is saying A is a list. B equals to

A. Now, remember what that does? That is

the assignment of we're saying B is the

same reference as A. That's what this

does here. The same reference. So does

it make sense to us that when we ask now

A is B. This should be true. And it is

like this is true because um

this is true because they are literally

the same reference. We're setting B

equals to A. So they are referring to

the same data. Now

um in terms of a variable reference they

so now their ids are the same

essentially their memory locations are

the same.

If you want to compare only data then uh

what operators have we looked at that

are comparison

we should think about that I mean we

just saw if we go back a couple slides

we have a bunch of operators to compare

data that's these guys right like equals

equals greater than less than so what we

could do is say does the data equal the

other data which would be something like

this equals equals operator

when two values are equal not the

reference ES. Does that make sense? This

equals equals is checking if two values

are the same, which is the the data, not

the not the reference.

Okay.

Now I want to show you a really powerful

um I want to show you a really powerful

miscellaneous operator which is going to

be uh which is going to be the in or

what's called the membership operator.

So this is an operator that checks if a

value is a member of a collection.

So this could be like uh this could be

like um you know where we have a list, a

tupil, a dictionary, just a collection

of data and we want to know is a value a

member of that collection which is

really useful for testing you know do we

have membership of something inside of

something else. So for instance, let's

say we have a list and we have a list A

equals and then we have 20 45 and 10

inside that list. So if we ask the

question

10 in A, this actually would return true

because 10 is a member of A. 10 is a

member of that list. So that is true.

Now that's useful to know. So the N

operator is a really powerful uh really

powerful operator.

Um, same same thing with not. So, we can

use not in. So, 10 not NA would be false

because there it is. We know it's a

member of A. So, that would be false. Of

course, it's NA. We can see it right

there. It's a member of that list.

Um, but if we check the different value

that's completely not inside of a, 30,

not NA, that would return true because

30 is not a member of that list. And by

the way, this in operator works for all

kinds of collections. So it would work

for a tupole, it would work for a set,

it would work for a dictionary.

Um, it would work it would work for all

kinds of collections.

Yeah, Roberto. Um, it's the fact that

there are different variables. So we may

sometimes it makes sense to have

different variables that are that maybe

they have the same value but they're

different variables altogether different

references

and the reason that is is because maybe

we have a copy of that data and then

maybe we manipulate the this one. Maybe

this is a copy of it and we manipulate

this guy and we leave this guy the same

to check the differences later.

Yes, references are tied to the variable

like A is a reference, B is a reference

even though their data that they're

pointing to is the same value.

Maybe we just have a copy of it that and

we manipulate one of those copies.

Yeah, that's why

What is the reason to check for what? If

they're equal, like as references, A is

B.

Honestly, there's not many good reasons.

Um, maybe if you want to know if

something is a copy of something else,

like you want to know that, like let's

say you're checking later down the code

and you have an A and a B and you want

to know if one is a copy of the other,

you can check and see if they're the

same reference.

That's the only reason I could think of

why you would do that. It's rarely used.

Rarely used, but it is an operator that

I wanted to show you in case you do

stumble across it um somewhere and

you're reading about Python or something

and you see the is

Like that's the only reason I can really

think of is to check if something is a

copy of another meaning it's the same

like maybe it's it's uh the same

reference

B equals to A then we could check A is B

and we know that then they're the same

reference.

Yes, the is operator is comparing

references not exact not the values.

Yes, that's true.

All right. How do we feel about this in

operator? Like the the membership

operator. Does that make sense? If

you're checking if an if a value is a

member of a collection,

that's going to be highly useful later

on. Highly useful. This one we will use

quite a bit. The is we will probably

rarely ever use, but but this one we

will definitely use.

Okay. So, I wanted to quiz you guys. Um,

what is the main difference between the

equals equals and the is operator? What

is the main difference?

We were we've just been discussing this,

so hopefully this this one is easy. Been

discussing it quite a bit.

Yeah, Roberto, now now you know the

answer. Perfect. Yeah, it is B. All you

guys answering B. Perfect. It is B. Good

job. You guys are right on top of that.

Good job. So, just wanted to point that

out. Like equals equals compares the

values. We're doing a comparison. Um is

checks if the references are the same,

which is the variable reference like the

memory location.

Yes, it is. Identity is the same as Yep,

that's what we mean. The references are

the same.

All right. So, wanted to do a short demo

on the operators on comparison, etc.

Wanted to just show off that demo um so

you can see and practice with it a

little bit. Uh so, let's hop over to

demo six inside of the lesson one. I'm

going to hop over to that.

which will be our last thing we will do

in lesson one and we'll move on to

lesson two.

Um, let me pull up demo six here.

Give me a moment.

All

right,

let me share my screen.

All right, so demo six. Uh, hopefully

you guys have access to this. This is

the last one inside of lesson one. Um,

now again, this one says try to use VS

Code. You can if you want. Again, Collab

works fine. You can use whatever you've

been using, Jupiter, Collab, VS Code,

whatever works for you. No big deal on

which one you use.

So, no worries on any of this.

Oh, yeah. So, um, F is So, yeah, this

that's a good question. What is F? So, F

tells Python to format. F is short for

format. It's basically format the the

string which we're going to print by

having some placeholders.

And um

we have variables called A and B. And

this fills in the blank of these

placeholders with whatever the values of

A and B are. So F just allows us to

format and fill in the blanks. Does that

make sense? Like wherever these um

braces are, we have a variable name

inside of it A and B. And we are um

going to fill in the blanks of those A

and B um by uh you know by just

replacing them whenever we do the print

function.

So there think of it as F is short for

format.

and we have a couple placeholders and

those will be filled in by our variables

A and B.

Okay, so this demo uh does a bunch of

operators. So it's going to do a bunch

of comparisons where we input one number

and we turn that into an integer. Input

another number, turn that into an

integer, and then do a bunch of

comparisons. So I'm going to jump over

to the notebook. I'll do that for us to

to show that off. But that's all we're

doing in in really in the beginning of

this uh demo. Um so let me jump over to

the notebook and show off that so we can

see it

but uh should be straightforward to

follow because we've done a lot of that

already.

Okay. So hopping over to notebook. Again

feel free to use whatever you want to

use. You can use collab, you can use um

uh Jupiter, you can use VS Code,

whatever you use. Um let's store a

variable as a and let's make it an

integer

and let's do enter your first number.

So we will do that.

Let's run that. So let's enter our first

number. Let's put in 10 or whatever you

want really, but I'm going to put in 10.

So that gets stored as a. Now let's do a

second number and let's do int

input um enter your second number

and let's input that.

So now I'm going to put in a second

number. Let's do 20.

So now we have a and b. Now let's do

some comparisons. So let's do print.

Um then we can do uh let's do the f

formatting like they had in there. Now

what this is going to be is we are going

to check a

um greater than or let's do yeah let's

do greater than b

um is

and then let's do uh comma a greater

than b.

Let's compare those two numbers. So now

this is doing the comparison. A greater

than b is going to compare those two

integers. What should this return? What

should a greater than b return? Should

it be true or should it be false?

Should be false. Right? So we should we

should uh display false here.

And that's what it is. 10 greater than

20 is false.

Okay, so that is false.

Let's do another one.

Let's do um let's try equals equals. So

let's say um a

equals equals to b is now what do we

think this one's going to be?

a equals equals to b.

What should this one be?

Yes, very good. This one should also be

false.

Let's run that.

And that one will be false. Very good.

Okay.

So, I think we get that. Let me give you

guys a Let me uh show you something

interesting. Let me uh go a little off

script from that demo document and let's

introduce a third number called C. Let's

do a third integer.

So enter your third number.

Let's do a third number.

Let's enter um another value of 10.

Okay. So, we have another number of 10.

Now, what I want to do is let's just do

uh multiple comparisons and do a logical

operator between them. So let's do um a

not equal to b

and

a less than c

or sorry b

less than c.

What do we think this is going to

return? This might be a little

challenging. What do you think this is

going to return?

We should use F. We can. I'm just not

printing. I'm just going to I'm just

going to run the cell. I'm I'm kind of

doing a shortcut and just print. I'm not

going to print. I'm just going to run

the cell and it should display what it

is.

Do we think it's going to be false?

Yeah, you guys are right on top of it.

Should be false. Now, let's break that

down. Why is that false? So, A not equal

to B is checking if A is not equal to B,

which is true.

A has a value of 10.

So, A has a value of 10. So, 10 is not

equal to 20. That makes sense. But 20 is

not less than 10. So this part is false.

So let's make a comment.

So the reason reason this is false

is because B is not less than C. So and

returns false.

Yes. Very good.

Now,

what if I take this code

and do this?

What's this?

Perfect. Yeah, you guys are right on top

of it. This should be true. And it is.

Now that's because the moment we have at

least one true which is going to be this

that makes sense right that is going to

be true.

Okay, one more and then we can wrap up

this demo. So I want to create a list.

I'm going to call it X and I'm going to

create a list of three numbers 10 20 30.

Okay. Now, what do you think uh this

result is?

What what should this be?

It should be true.

Very good. Should be true.

Now what should what should this be?

This is also true. Very good. This

should this should be true because this

is not going to be inside of the list.

So that makes sense. That is not true.

Now what I want you to notice is nothing

will change if I change this to a

tupole.

Nothing will change. We can still check.

So we can still check if uh 10 is a

member of this tupole and we can still

check if 40 is not a member of this

tupole. So nothing really changes,

right? It's still in membership operator

in checks is it a member of any

collection?

Okay,

very good. Any questions about the these

examples?

Any questions? Do we feel comfortable

with some of these operators? You guys

were right on top of it. It was very

impressive. You guys got those right

away.

Uh, you change print

a= b and a is b. And I entered four both

times

and I got the same result. True.

Uh so you had your A and your B. So you

you had um

you had A equals to 4

and B equals to 4 and you checked

uh you checked this.

Yes.

Yeah. So that is a little confusing and

I can understand why. So the reason this

ends up being true is because this is a

scalar data. So for scalar data it's

going to optimize in the memory to point

because four the integer four occupies

the same memory address always. But when

we create an array, when we create a

list, that is um a new object.

So yeah, that one's a little confusing,

but it's it's only because this is

scalar data that um Python kind of

knows, okay, four is the same like

integer in the me in memory always

regardless of if we're referencing it

from this from two different variables.

That's that's the reason I um yeah I

forgot to mention that example but it's

purely so the the reason this is true is

because

this is true because of scalar

data optimization. Essentially it's it's

not going to waste creating a new object

when it's just a single integer four. it

basically occupies the same memory

address

uh as as a as a reference.

But when we build a list like that is a

different object.

Okay.

Very good.

All right. So,

um, that will wrap up lesson one. What

I'm going to do is go into lesson two.

I'm going to pull up the lesson two

notes, the slides for lesson two. So, if

you have those, let's pull those up. Um,

I do encourage you now, there is a

guided practice at the end of lesson

one, and that is for your yourself. I

would encourage you as as kind of

homework between now and and the next

time we meet um to do the guided

practice for lesson one. Try that out on

your own. It is um you guys should have

access to it from your LMS and the

reference materials. There's a guided

practice. Try that. Try those out. Okay?

Try out the guided practice for lesson

one. It's just it's just some additional

practice of the things we just covered.

Okay?

Let me open up

um

lesson two.

Give me a moment here.

Okay.

So, let me share my screen.

Okay, so now we're going to move on to

looking more closely at those things

like lists, tupils, dictionaries, and

then looking at control flow with

conditional statements and loops. So

we're going to get into the fun stuff, I

would think, um that you guys may may

have been waiting for.

Okay, so we've talked about uh we just

finished talking about lesson one where

we have you know Python as a really

important thing to learn and study

because it's used all over the place

with data science and a IML. So one of

the things is we need to continue

learning about it with things like loops

things like if else statements to

control the flow of our programs and

these basic data structures.

Let's continue forward. Um so this

lesson we're going to talk about list

tupils dictionary sets uh we're going to

you know talk about the differences how

we can access data within things like

list how we can modify them um how we

can access things from tupils what are

the differences we'll review all that

one of the big topics is going to be to

um look at how we can control the flow

meaning control the flow is using like

decision logic like if this is true then

do this else do this we'll talk about

those kind of statements. We'll talk

about iteration. So how we can do loops

to repeat um statements of code that we

want to do. Um we'll also talk about

organizing our code a bit into

functions. Um which is going to be

really helpful to for our own

organization and reuse and

maintainability.

Okay.

So let's jump into it. Let's talk about

some of those data structures.

So, we've already talked about that

these aggregate data structures exist um

and they allow us to manipulate data

inside of Python. Lists, tupil, sets,

dictionaries are the main ones we're

going to focus on.

Let's start with lists. And I think

lists are going to be something we're

going to use quite a bit of throughout.

So, they're going to be a really good

place to start with. They are super

popular in Python. A lot of people um

use them to do to work with data. They

are kind of the most basic um data most

basic and useful data structure that

there there is.

So what is a list? A list is a an

ordered mutable meaning it can be

modified data structure that can hold

elements of different data types. So

it's a collection of data and there's no

requirement that all the members of the

list be the same type. In fact, you can

have different types. You can have

integers, you can have floats, you can

have strings,

um you can have even more abstract

objects be members of a list. Um so any

kind of data can live within a list. But

the big thing is that it is modifiable.

It's dynamic. You can change you can add

things to it. You can remove you can

change things. Um and it has an inherent

ordering which is nice. So you can you

can be reassured that there is some

inherent position of items and it will

maintain that order. So we can access

things based on the order like we can

access the first can access the last we

can access anything in between.

So that's nice. Um so what are some key

characteristics? So uh lists support

multiple data types. We talked about

that. There's no requirement that

they're all the same. They can have

multiple. They allow for indexing, which

we're going to talk about. This means

that we can access things based on their

index, which is their position within

the list. So, we can always access the

first thing, the last thing, anything in

between, based on its position. That's

another word for uh index because lists

have natural ordering to them which is

um really really powerful to ensure that

there's um one you know there's a first

position a second position third etc. So

lists are really nice for that.

Um

uh they are modifiable which is really

nice. So we can add things into the

list. It's very dynamic. So once we

create a list, we can throughout our

program, we can add data to it, we can

remove it, we can change. Um, lists also

allow duplicates, which may be

desirable. Like maybe we add something

into our list that already exists.

That's okay. List allow duplicates. This

is going to be different than sets. Sets

do not allow duplicates. Sets are just a

bucket of unique things. So if we added

a duplicate into a set, it would reject

it. we wouldn't have any errors, but it

just wouldn't um show up as a copy. We

would just have a set is only going to

maintain one copy of an item. It it only

allows unique items. Lists allow you can

have as many copies of data as you want

inside of a list. Um so it does allow

duplicates.

Now, we already saw in terms of syntax,

lists are um defined by brackets. So

when you see those brackets um it

defines a list and its items are

separated by uh commas.

So

we can have a list that looks like this.

Yeah, I was going to explain slice. We

have a a couple slides about slicing

coming up, but slicing just means that

we can slicing means that we can grab a

section of elements at a time from a

list. So for instance we can uh actually

let me use this example down here. A

slice would mean we can grab like these

first three slice of the of the list or

we can grab the last five elements or

whatever like this is a slice. It's just

a a subset of the list that we can grab

we can access.

So slicing just means taking a subset,

taking a smaller section of the list and

we can grab all those elements at a

time. And what that's called is within a

slice.

Yeah, like a slice of pizza. We're

taking the whole thing and we're taking

a small section of it.

So that's and actually that's going to

be possible because the list has a

natural order to it.

Does the data have to be sequential? No,

it doesn't have to be. In fact, you it

can be completely different types.

Does that make sense? Like so look at

this example down here. Like we have 10

2 5 hello. That's a valid list. You can

have different types of data in there

which isn't sequential at all

in a slice. No, it doesn't have to be.

So you can have like you can have a

slice that picks every third element

uh every other element. Um yeah, it

doesn't have to be sequential. No, it

can be customizable.

What's also nice is you can slice from

the beginning or you can slice from the

end as well. So you can you can go from

the end and slice backwards. Um or you

can go from the beginning and slice

forwards. So you can grab like every

other element from the beginning. You

can grab every other element from the

back and work your way forward and stop

at a certain point.

Slicing is very nice. Yeah. So, I'm

going to show us how to do that.

Can you slice in the middle? Yeah, you

can slice anywhere you want.

Can you slice a pizza in the middle?

Sure. Would you do that? Maybe not. But

yeah, you can slice anywhere in the

list. You can slice.

Does slicing change the original list?

No, it's just selection of a subset.

No, it just it just extracts elements.

It doesn't like permanently change it in

any way. It just gives you a view. Think

of it as like giving you a view of that

subset.

Okay, so you may be wondering when would

I ever use list? So normally you use

lists whenever you want an ordered

collection that is dynamic meaning it

may you may want to add things from it.

Um we may want to modify things from it

frequently. Um but we want something

that is dynamic and has an ordering to

it. Lists are perfect for that reason.

So they can contain data that we can add

to remove from change. So lists are very

versatile. I think most use cases with

manipulating

um a collection of items would fall

under a list. A list would be a very

good choice. Um

we can slice elements and use them.

Yeah. Yeah. You can slice you can store

a slice inside of a variable and use use

the use that resulting variable. Yeah,

for sure.

order information like

Yeah. Yeah, that's a good example. Yeah,

order like the collection of orders

would probably be in a list because it

can change. Um, but the prices would be

probably static. So, that would be more

uh suitable for a tupole. Yeah, a tupole

or maybe a dictionary. A dictionary is

probably better because you can have

like a product ID or a name that maps to

a price. Probably a dictionary would be

more appropriate for a price list. But

but yeah, tupole maybe makes sense too.

By the way, look at this list. Do you

guys see how it has different types of

data in it, right? It has so like the

first element is an integer, the next is

a string, the next is a float, the last

is a boolean. That's totally valid in

Python, which is kind of unique to

Python. Like a lot of languages don't

support that. A mixtyped array.

Don't really they don't really have that

notion of that.

All right. So, what I wanted to get to

was the positions. So, this is going to

be this is really important to pay

attention to because this is going to be

something that we will be using

throughout the program is how to access

data by its position.

Okay, which so there's another word for

that. The position um sometimes you will

hear called the index. So the index in

the list of where these where data

members live is their order like their

position amongst amongst the list.

So what's special about Python is that

it has the first position

is index zero which trips people up all

the time. The very classic trip up of

Python is that the first element of a

list is at position zero. The next

element is at position one. The next

element is at position two. On and on

and on. The last element is at position

n minus one where n is the size of the

um size of list. So, however many

elements we have in our list.

Um,

so

if you want to access the first element,

you would be looking for the element

that's at position zero. If you want to

access the second element, that's this

this name Bob, that is at position

number one or index one. So that's

something that trips up people is that

it actually starts

starts at zero

which is really important to understand

is that the positions start at zero.

Why is that? Um that's a good question.

I mean

different so so different languages

treat that differently. Um,

it's more historical reasons that it was

created that way. Uh, but I think it's

based on your like

think of it as like how far into the

list you are. So, if you're in the

beginning, that means you're basically

at at um position zero because you

haven't made any progress like

traversing the list. I think that was

the intuition.

Oh, yeah. which is kind of what Tim says

there is like yeah if you're at the very

beginning of the list you're at position

zero because you haven't made any you

haven't made any forward progress in

traversing so it's like you're at step

zero you're at the beginning

okay but okay so aside from why it was

that way does it make sense that that

the first like what I'm saying is the

position zero is the first element,

position one is the second element,

position two is the third element, and

on and on and on.

That's that's how it works in Python.

Okay.

Now,

what I'm telling you is the positions if

you were to view it going left to right.

So, in other words, in the forward

direction, we start from zero and go up

to the to the n minus one in terms of

the position. Now, what's really

convenient is the positions can also be

indexed from back to front, meaning they

can also be indexed going this way.

And what's really nice is the very last

element starts at position minus one.

The very last element on the right

starts at minus1 and then goes all the

way up to minus n.

Okay, now that's really convenient

because if we want to access the very

last element, I don't need to know how

big the list is. I just need to access

the position minus one. That guarantees

to access the very last element. The

minus one position is the very last

element.

And then it it so this is the very last

element. This is the second to last,

third to last, fourth to last, on and on

and on. And then by the time you get to

the front, it is minus n, which is like

the

uh number of elements in the list minus.

So in this case, minus 6 is the very

beginning.

Now, why why in the world do we care

about negative position? It's for that

exact reason.

we can um we can think about the

positions from the end of the list going

backward which is very powerful like if

I want to access things from the very

end I don't need to know exactly how

many there are I just need to know minus

one is the end minus two is the second

to last minus3 is the third from last um

which is pretty convenient

okay

so does that make sense on the negative

index index. It's the negative index is

going from right to left. It's from the

back to the front. Minus one is the last

element.

And then you go second to last, third to

last, right? And that is minus 2, -3,

-4.

So from right to left. Yeah. Right.

Right to left. minus 6 is actually the

first element

because it's six back from the end which

is the f which is the front. There's

only six elements.

Yeah.

So are there any question about this is

really important to understand really

important to understand because this is

how we are going to slice and access

data is based on these positions.

Uh is there syntax to get the number of

elements in a list? Yes, it's the length

function which is len. So length of list

would give you uh the number of elements

which in this case is six. So ln the

length function gives you how many

elements are in the list.

len or length. So this guy this

function.

Uh when are you counting backwards?

You're usually counting backwards when

you want to know what's at the end of

the list. So when you only care what has

been added at the end and you want to

maybe you want to get the last five

elements.

So you slice backwards from the from the

end.

That's that may be useful like maybe you

want to know like it imagine a list is

holding in your orders

and so you want the last five orders. So

you can just go from the back and go

towards the front. Min -1, -2, -3, -4,

-5

would be the last five orders. There's

going to be many scenarios where we want

to count from the end.

uh when we're manipulating data

later on when we're when we're working

with um bigger sets of data um in

something like an like a matrix, it's

going to be useful to grab like the last

five rows, last 10 rows.

So going from the end makes more sense.

So you imagine if we have like a larger

matrix of data um maybe we want to slice

out these last 10 rows in which case we

want to count from the minus like the

last row is minus one and we want to go

back towards the front.

Can you sort? Yes, you can sort. Uh I

I'll have an example of that in a

second. Yeah, you can sort.

There's a built-in sorting function in

Python that that will allow you to sort

data in a list. Yes,

there's a built-in function for that.

Okay. What I wanted to do is show you

guys an example of accessing elements

from the list. So assuming we know the

position which is that index all we have

to do is use brackets to access items of

a list. So imagine we have this list

called fruits which has some strings in

it apple banana cherry mango and we want

to access the first element. So that is

just this code here fruits bracket zero.

So the bracket tells the interpreter,

hey, I want to access something within

this list. And then all we have to do is

give it the position that we want to

access. So this is position zero, which

is going to be the first element. This

would retrieve the first element. Now,

it's not removing it. It's just

accessing it so we can view what that

is. So, it's actually going to give us a

copy of what that value is under the

hood. Basically, a copy of it. It's not

going to permanently. It's not going to

delete it. It's not going to remove it.

It's going to give us a

copy of what is at position zero. In

this case, apple.

Now if we put in a position two in that

bracket that should be remember the

indexing is 0 1 2 3

because there's four items. So the last

element is n minus one which is three.

So the item that's at position two is

going to be cherry. So this should

return the string cherry.

Okay. So we use we always use this kind

of syntax

position

to access the element that is at that

position.

Okay, pretty simple. And what's really

nice by the way about this is this the

same exact thing works for tupils

because tupils also are ordered. So

we're going to see that when we get into

tupils. But the same exact things works

where a tupole has a first item, a

second item, a third which are index

zero, 1, two, three. And we can access

things in the exact same way. So we can

access a tupole by its position as well.

Same exact thing will happen position.

So we can get the first item of a tupole

with with uh by doing um zero and we can

get the the last by doing minus one.

and on and on.

Any questions about the accessing?

Can we get position based on value? Yes.

So that is a special function called

index.

So if you do like list dot dot index

and then you pass in a value like apple,

this would return to you the index of

the first occurrence. Not every

occurrence, but just the first. So like

because we could have multiple copies of

Apple in our list, but the first time we

stumble upon Apple, this would return

this would return uh zero because apple

occurs at index zero.

So index function

uh would return um the index

which is which is the same word as

position.

Okay.

All right. Any other questions? We can

uh I think what we'll do is we can take

if unless there's any other questions we

can take a short break and then come

back and continue.

Uh, I want to talk about slicing.

No, it would return none. It would

return null. Basically none

as opening closing braces are square on

it if it's a list. No, it's actually the

same as a tupole. Uh in terms of

accessing it's the same

uh

it's it's in terms of accessing it's the

same. Um, for creating a list, yes, it

is square brackets. This creates a list.

The brackets always create a list. For

accessing elements, this is the same as

a tupole. The the square brackets for

accessing.

Actually, in most things in Python, it's

the same where we uh access data using

the square brackets. A dictionary, same

thing. We access keys by using the

square brackets.

So square brackets in Python is is

basically like an access operator.

Uh no, so list.index doesn't only work

for the first item. What I'm saying is

you can have lists that have multiple

I'm saying it returns to you the first

occurrence,

the index of the first occurrence

because we could have a copy of Apple

later on in the list, right? So there's

nothing that stops us from having a

duplicate. So, we could have another

apple down here. Like, let's say we had

another apple at the end of the list. If

I did list.index apple, it's only going

to return to me this first one, even

though there's another copy of it in the

list.

So, it only returns to you the first

occurrence.

But, you know, there's nothing stopping

me from doing list.index of banana or

cherry or whatever.

Are there data types by this for all

Unicode characters? Um, I think there's

Yeah, I think there's special strings

you can make that do Unicode,

but I'm honestly not 100% sure.

I would research that. I don't really

know. I think you can do that with

strings,

special strings.

Yeah, I don't think there's anything

like I don't think there's anything

inherently special about it that Python

can't handle. It's just you sometimes

have to like escape characters

uh

to distinguish them. But

yeah.

All right. Going back to our example,

um we had the uh going back to that

fruits list.

We can access things from the end. So

here's an example of us using the

negative index, right? So minus one

grabs the element that's at the last uh

that's the last member of the list. So

mango minus3

um would be the uh third from last,

which is going to be banana.

So negative index really handy to access

from the end of the list going backward.

Um pretty useful there.

There's an example of it.

All right.

All right. Let's talk about slicing.

Okay. Let's talk about slicing. So,

slicing allows us to extract a subset of

items from a list using a specific range

of indices. And so the syntax to do

slicing

is going to be using our brackets again,

but um we use a colon to specify where

we are starting and stopping our

positions. And not only that, but we

also use uh a colon to signal how many

we want to step by. So do we want to do

every other in which case we would step

by two. Do you want to do every third

element which would step by three? So

most slicing is going to follow um this

sort of syntax where we do a list and

then we do like a start index and then

we do colon

and then we do stop index

um and then we do colon uh step. Now

what you will see is that the the step

defaults to a step size of one meaning

we grab every element in between

starting and stop. Um so step size of

one is the default. So we actually

typically will not include the step size

unless we specifically want to get every

other which would be a step size of two

or every third or every fourth or every

fifth. Um so usually we leave off this

step size and we just we just have a

starting and a stop as part of our

slice. Um and what what that does is it

tells the interpreter to access

everything between this start and stop.

So um with one catch which is a very

important catch that trips everyone up

which is that Python um is very annoying

and that it uh when you do slicing it

allows you to include the starting

index. So uh if we start somewhere we

will guarantee that the slice will

include that but it will not include the

stopping index. it will do everything

between there up to the stop index but

not actually including what's what's at

the stop position.

So for example,

this slice that you see um on the screen

is a slice that would be starting at

position two

because remember um let me draw this

out. This is position zero. This is

position one, position two, position

three, position four, five, and six. And

so this is a slice that would start at

position two

is our start.

And meaning we're guaranteed to get 34

because we're starting there. But the

stop for this slice would have to be

here. This is our stop because we are

going to include everything in between

there. We're going to include all of

this this uh slice as part of our uh

what we can access. Um so that would be

positions 2, three, and four. So this

slice would be um basically like this

list. Um it would start at two and go to

five.

And then it technically would be a step

size of one, but remember we don't

really need to include that. So this

would really be list um two to five.

That does feel annoying. Yes, it's it's

because they you have to know where to

start and stop. And so they cho Python

chooses to to be um not inclusive of the

stopping index, but it it includes the

starting index. Um and and that's just a

choice. That's a design choice of

Python.

Okay. So the colon gives us the slice.

Um, so if you see if you see uh uh if

you see

a colon inside of a brackets, that

signals you're grabbing a collection of

items. So we So this this actually

returns a smaller list. This returns a

list of 34, 20, and 80. So it's a it's a

slice meaning we get multiple items

rather than just a single item from from

the selection.

No, 54 would not make the cut. 67

doesn't make the cut either because

remember we don't include the stopping

index.

Can you do two to four plus one? Yes,

you can do arithmetic in there and

Python will evaluate the arithmetic

first. So it will do 4 + 1 first and

determine that is five.

Yes, you can do that.

Okay. So before

before I move on, uh does the slicing

idea make sense? It is grabbing a

collection of items from the larger list

and we set it up with this syntax.

Uh what do you think? What do you think

0 to six would return? What would be

your guess?

What's included in the slice though?

It's not just 67. If you go 0 to six,

how many? Like, you should be getting

more than that. Yeah, you should get

everything, right? Exactly. You should

get all of those numbers up to

uh 54. You would not include 54. So you

would get 76, you get 12, you get 34,

20, 80, and 67.

Yes, that's true.

Yeah. Yeah. So the the step relevance is

that we can so from our slice we can

choose um our step size of how many

element like how much we want to skip

positions within that slice. So step of

one which is a default means we get

every position we go we increment by one

one position to the next to the next to

the next. A step of two

uh

a step of two would be that uh I'm going

to grab every other element from the

start. So a step of two like if I sliced

this

and changed this to a step size of two,

then that would only grab this and this

because that would step over. It would

take two steps to get to the next

element of the slice.

Whereas a step size of one is going to

grab everything

because it's going to go one index to

the next. So think of the step size as

how many positions are we incrementing?

Yes. 2 to 7 would include 54. Yes,

that's right.

2 to 7 would include 54.

Okay.

All right. Let's see some more examples.

So if we have a list like this and we

slice it from 1 to 4, the output is

going to start at index one

and go all the way up to index 4 but not

include four. So index 4 is this guy and

it's not going to include that. So it's

going to be these three elements here

would be our slice. Yep. 20 30 40. Very

good. That's what it would be.

So that's pretty useful

to be able to slice. Let me show you

another example.

Oops, don't have another. Let me go

back.

Let me show you another example with

this. Um, so what we can do is we can

actually slice backwards as well. So,

what do you Let me ask you guys this.

What do you think this slice would be?

Actually, let me erase this. Let me do

minus

or

what do you think this would return?

We can use negative index index uh

indices in our slices.

What do you think that would be?

Very good. 30 40 50. So it's going to

go. So remember -4 is the fourth from

the last. So it's going to be this is

minus4

and then this is minus one. So we're not

going to include minus one. So we should

be doing this slice here.

That should be the slice. 30 40 50 would

be that.

Okay. One other a couple other examples

I wanted to give you is that you can act

in certain special cases you can

actually leave off the starting and

stop. And what that would signal the

interpreter is that you want to go all

the way to the end or start all the way

from the beginning. So if you do

something like this,

let me show you an example. If you do

something like this and you do not

include, you leave the the start blank.

You leave the start blank and you go all

the way up to minus one. What that would

signal to the interpreter is by default

um start at the beginning. So start at

zero. Essentially start at zero. If you

leave off a slice

uh as your start that the interpreter

assumes you want to start at the

beginning.

So what do you think this slice would be

knowing that?

Yes, exactly. You guys you guys are

right on top of it. 10 to 50. Perfect.

So it's going to be everything but the

last

everything but the last would be

included in that slice.

Perfect. And so the other thing is we

can leave off the end which would signal

that we want to go all the way to the

end. So what do you guys think this is

going to be?

What would that be?

Yes. So this is So this is actually

going to include the end. So I know

that's a little counterintuitive, but

it's actually this this guarantees we

include the end. So if it's blank, it's

going to go all the way to the end,

including the end. I know that's that's

annoying. I don't blame you for thinking

it should be 50,

but it basically goes to n. Basically

goes to n, which would mean that we

remember the index is n minus one is the

last index.

So, so if it's blank, this means we

should go we should end at n

which is the length of the list meaning

that um the last element is at n minus

one. So we should include the n minus

one. So yeah that would be 20 to 60

actually sorry 30 to 60 because we uh

index two is 30. So that would be So

that slice would be this one all the way

to the end.

Good. I have one more example for you

that's really going to throw you for a

loop is this one. So what happens if we

have

this case?

This probably won't for loop, but I'll

have one more following that.

All yeah, 10 to 60 everything. Perfect.

You guys are right on top of that

because we're leaving the starting blank

meaning that we should start at the

front. We're leaving the end blank

meaning we should go all the way to the

end. So that would be everything

everything in between. So at that point

we're not really slicing anything,

right? We're not really slicing much.

We're just taking the whole list.

Now, one example I want to give you

is

what if we sliced

and we had a step size

of minus one.

Any ideas what that would do?

What does this do?

What is a step size of minus one?

So minus negative index goes from the

back, right?

So if we're step sizing minus one, what

should we be doing effectively?

So, so this is signaling we basically

want to have the whole list but step by

minus one.

Yeah. So this this would be the reverse.

So this would be the reverse list. So, I

know that seems wacky, but that actually

is a way to validly reverse a list in

Python is to index to step size by minus

one. Because what that means is you're

slicing the entire array, but you're

stepping by minus one. We know minus one

um we know minus one goes

backwards, right? Effectively, because

it starts from the end. So, step size by

minus one would be go back this way.

each element going back this way. So

that would that would effectively

reverse the list.

Yeah, there now there is a reverse

function. A list has a reverse function.

So that is in English. But this is like

this is an alternative to to reversing

minus one

step size of minus one.

Uh, no, Roberto, they're not quite the

same because remember when you do two

colon and then you leave out the blank,

that means you're you're going all the

way to the end, including the blank,

including the end.

So, -4 to minus one would be the same as

2 to

uh 2 to six or 2 to 7. Two to six.

Sorry. Two to six.

60 to 10. Yep. It would be it. So this

reverses it. Meaning this would be this

would return to 60 then 50 then 40. It's

the reverse of the list when you step

size by minus one.

I am sure.

How about slice from position one?

Slice from position one to second to

last and reverse it.

Uh what do you think that would be?

You're starting at one going to second

to last.

And then rever like we already know

what's reversing is step size minus one

that will always reverse.

What is colon colon?

Colon is is the fact that we're leaving

colon colon is not anything special.

It's the fact that we're leaving the

starting. It's just the syntax of

slicing, right? colon colon is because

we are we're leaving the starting and

stopping blank

which we can do. We're allowed to do in

slicing. So that would signal that we're

doing everything but we're going step

size minus one.

Yeah.

Okay.

Any

other any other questions?

How do we feel about slicing? Do you

feel okay with it? Are we going to we're

going to practice it more as we go

along? Uh because we're going to use

slicing quite a bit when we work with

data, but does the concept of slicing

make sense? Yeah, you need you need more

p We'll do more of it. We're going to do

slicing throughout the program.

Yes, we'll do multi-dimensional uh in

our next course.

Not right now, but in our in our data

science course, we'll do

multi-dimensional.

that help?

Yeah. Step can be a very Yeah, sure.

Sure. So, there's there's nothing that's

stopping you from, you know, there's

nothing that's stopping you from doing

like let's say x is two and then we do

um numbers

numbers and then we have uh two to to

six and then x

Yeah, that's fine. There's nothing that

would stop us from doing that.

I do you mean that I think that's I

think that's fine. There there would be

nothing wrong with that.

But you're right that it should be an

integer. If it's if it's like a float,

Python will complain. It needs to be an

integer step size and it needs to be it

needs to be uh

um in order to get any meaningful data.

We wouldn't want that step size to be

too big or like it, you know, if we pick

it to be like 20 and there's only five

elements, that's not going to make any

sense. The step size needs to be

reasonable.

All right. So, what I want to do is show

you some functions that lists have. So

probably one of the most useful

functions a list has is the ability to

append items to the list. Now this will

add items to the list and particularly

it will add it at the end. So this is

this is something we will use quite a

bit is the list.append

function.

So this will uh this will add this will

modify the list and add a new element at

the end. So append always appends to the

end. Um

and so this will uh allow us to um take

this this string cherry and now the when

we append it this list is now

permanently been changed to have cherry

at the very end. So there it is. It is

now at the back of the list and it is it

is now at the kind of end position when

we do append. So here you see the list

being really dynamic allowing us to add

elements to it through this append

function.

So append really really useful allows us

to to add we just pass in we pass in an

element inside of the append uh function

here that we want to add to the list and

it will it will go to the back of the

list.

Um, lists also have a pop function

um, which you pass in a position and it

will remove that item that's at that

position. And not only will it

permanently remove the item that's at

that position, but it will return it

back to you. So, pop is really useful if

you want to remove things um from the

from the list. um if you don't provide a

position. So if you don't provide any

index that you want to remove from and

you just do if you if you just do um pop

without any uh thing in there that will

always remove the last element by

default. So always just so if we just

did this it would remove the 40.

Can we append in a specific position?

Um, yes. You would use the insert

function and then give it the index you

want to insert into. So,

would would uh be every list has ainsert

function to to and then you put in a

position you want to add it to.

Is it common use for append and remove

during? Yes. So append is really common

to add new things to the list which

maybe we're doing like data aggregation.

We want to add things to a list and then

take the average of the list. That's

very common. Um remove. Yes. Maybe we're

working our way through a collection of

things and when we process it we want to

remove it. So we can do pop to remove it

from the list.

Yes.

Uh yes. So when so when we pop it

permanently affects the list. So um

everything gets shifted. Yes. All their

positions get shifted according to what

we removed. So like in this example um

30 is uh 30 is index um two but when we

pop it now becomes uh the last element.

So it would now be eligible to be index

minus one, right? Cuz when we remove 40,

30 is now the end of the list. Um,

for example,

so yeah, everything shifts

and you know this this example here um

pops from index two. So we would go to

index two, which is 30, and remove that.

And so 40 now shifts up to be at index 2

whereas previously it was at index 3.

Insert as well. Yep. When you insert

everything shifts. Yep.

Okay. Here's a here's a really useful

function as well. So we have the extend

function

um which would allow us to add in

multiple elements. So this is the same

as if we appended every individual item

in this collection to the list. So

extend takes a list and adds its

elements to the other list. So notice

that we have um

we have a list of colors here, red and

blue, and we're extending it with a list

of green and yellow, which will result

in the colors list now having all of

those elements. So extend is really

helpful if we want to add in multiple

pieces of data to an existing list.

a lot of questions. Um, can we get the

index and values with a print command?

Uh, yeah.

Yeah. I mean, you can use a I'm not sure

what example you have in mind, but yes.

Yes. Append is one item. Extend is is

taking an entire list and and adding all

of those elements to the existing list.

Yes. Append is for only one value at a

time. Yes. Extend is when you're adding

multiple values.

Append is one value at a time. Yes.

Uh okay. So let me ask you guys what do

you think of this? Um,

which of the following method adds a

single element at the end of the list?

So, adding a single element at the end

of the list.

Very good. It should be a Yeah, we

append. Append adds and append always

adds to the end.

Very good.

Okay. So, now we're going to have a

demo. Um, and by the way, this is um

this is a demo that uh is an existing

notebook. So, um, what you would want to

do is, especially if you're working in

collab, is take the notebook. Now, this

is within lesson two. So, we're going to

do demo one and lesson two. You would

want to take that notebook if you're

working in Collab and upload it. I'll

show you how to do that, but we're going

to do we're going to do the demo that's

inside of uh the first demo inside of

lesson two.

Um,

so let me share my screen.

Okay. So, if you're inside of Collab,

what you're going to want to do is go to

file and then upload notebook. So,

you're going to want to go to upload

notebook and then um pick the hopefully

you've downloaded the demos in which

case you have the notebook from from

lesson two. There's a bunch of notebook

files, the IP YMBs. You want to upload

um those demo those demo notebooks.

Okay. If you're working in collab, if

you're working in Jupiter, um, or you're

working in, uh, VS Code, you can just

open that file, uh, within VS Code or

Jupiter, um,

and, uh, you should be good to go from

there. So, I've I've already uh, done

that. This is this is demo one inside of

lesson two. Do you guys have access to

that notebook?

Demo one and lesson two.

There should be lesson two has a bunch

of notebooks that we're going to work

through. Um,

okay.

Very good.

So, if we run if we run this piece of

code um that's in this first set or

sorry first cell, it's going to um

create this list which has different mix

types. So this list has integers, it has

strings, it has floats, but we can

create this list. If we just hit run,

um, we now have a list. And what I want

to show you is if we were to check the

type of this my list, um, of course,

this should be a list, which it is.

Okay. Uh, thank you for uploading that.

Perfect.

Okay. So let's go ahead and access a few

elements. So we can access the first

element here. We can access the element

the fourth which would be at position

three and then the seventh which would

be at position six. We can access all of

those

and we put those into a new list here by

putting them inside of the brackets. So

that that means that we're accessing

this first one. That's the first element

of this list that we're creating. It's

25, which is here.

By the way, what do you guys think

happens if we try to if we try to use an

index that is too big for this list?

What do you think would happen? Like if

we if we tried to do if we tried to use

code that would be like um my list and

then we put in the index like 20. What

do you think would happen?

because there's definitely not 20 items

in this list.

Yeah, it'll be an error. So, let's try

running that. This will give me an error

that says it's out of range. Yeah, an

exception, right? It would be an

exception, which would say, uh, we have

an index error. Um, we're trying to use

an index that's too big for our list

essentially.

So, just pointing that out. Um, let me

make a comment there.

Um, this

uh index is out of range for our list.

So, we should get an error.

See how I'm making a comment? Making a

comment there to remind myself of why I

got this error. So, remember, comments

are useful.

All right. So, we access things and we

can uh put those inside of a list. Now,

let's do negative index. So, we know

negative -1 should give us the item

that's at the very end of the list,

which would be this 2.718.

Um, so that should be there. And then

minus 4 would be fourth from the back.

Minus 7 would be seventh from the back.

So, we can uh get those values. Not too

bad.

Then we have a slicing example.

So 2 to 7 we know as a slice. This

should um this is a slice that uh slice

that starts at index 2 and goes to index

7 but doesn't include index 7.

Right? So that should be the slice. Um,

so if we run this code, um, this would

extract everything starting at position

two, which should be the third item of

the list, all the way up to, uh,

position 7.

And one other thing I wanted to show you

is we can extract how many elements are

in the list.

So I wanted to show you guys that this

code tells us the length of the list

which would be if we did length of my

list.

Yeah. Len. So we pass that in the the

length function. We pass in my list um

which should give us 10. So there's 10

items in this list.

So just wanted to call out that there is

this length function that we can do with

the list.

Can you show printing index numbers for

the list?

Like do you mean an uh every number in

its index?

Do you mean that every number and its

index? Yeah. So the code that does that

is the enumerate function and we we

would use a loop. Um so it would be

something like for index

um value in enumerate

uh my list and then we could do um print

uh index

and value

like that

you know we haven't learned this We

haven't learned any of this yet, but

that's that's what it's doing.

Yeah.

Okay,

cool. Uh let's see. So um finally what I

wanted to show is that we can append. So

if we uh take our list this is what it

currently is

and then we append a new element we can

print out the list and you can see how

it ends up at the end. So we take that

original list and we just add a new

element at the end. Um and then this by

the way we didn't we didn't uh explain

this but remove will find that value

find the value 100 and remove it from

the list. It will find the first

occurrence.

Yeah. So remove will find the first

occurrence of this value. Pop is index

based. Remove is value based. So remove

will look for the 100 and and take that

out of the list. But pop will um be

index based. So if we do um

if we do so pop is index based. So we

could uh do my list.pop

pop and we could pass in a zero which

should remove um this will remove the

first element

and then we can uh print my list.

So that removed the 25.

Does remove all? No, I think it's just

the first occurrence.

This is the first occurrence. You'd have

to do it multiple times if you have

duplicates.

I think I have to double check that, but

I think it's just the first occurrence.

Okay. All right. So, I know we're a

minute over.

Uh, thank you guys so much. What a great

first couple of sessions. Um, I think

we're picking up this really well. So

very good job. A lot of great questions,

a lot of good um back and forth. So I

appreciate that. Hope you guys are

learning and picking up this Python as

we go along. Um we have a lot more to

cover. So uh you know next time we meet

um you know we will uh continue talking

about the other data structures. So we

have to talk about sets, tupils,

dictionaries and then we have to get

into uh loops and if else statements and

then we'll eventually work our way to

functions. So, a lot more to cover, but

we'll get there. Um, and but hopefully

you guys are learning a lot. Any

homework to do? Not formally, but I

would request that you guys work on the

guided practice for lesson one. Work on

the guided practice for lesson one if

you can.

Okay? So, go into your reference

materials, find the guided practices,

work on the lesson one guided practice

between now and our next session.

Thank you guys. Thank you so much. Uh

thank you for I know these 4hour

sessions are a lot. Appreciate your

patience. Thank you so much.

Have a great rest of your week.

>> So you might be wondering what's changed

in machine learning and why is it the

best time to get into it now. Well,

let's go back a few years. In the past,

machine learning was more about building

models based on historical data. It was

about training algorithms to predict

specific outcomes like classifying

emails, spam or not spam, predicting

house prices based on past data. But

fast forward to 2026 and the landscape

has changed dramatically. Today, machine

learning isn't just about making

predictions. It's about building systems

that can learn, adapt, and improve over

time. We're no longer creating

algorithms to just run experiments

offline. Now, machine learning systems

are integrated into real world

operations and are capable of making

decisions that impact the business

immediately. Here's an example to make

it clearer. In the past, an e-commerce

website might use machine learning to

predict what products a customer might

want based on their past purchases. Now,

the systems can constantly learn from

new customer data, continuously refining

those predictions in real time as

customers preferences are changing. The

world of machine learning has evolved

from theory to practice and this has

created a huge demand for machine

learning engineers who can build

scalable systems and make them work in

real world environments. The impact of

machine learning is now directly tied to

business outcomes and machine learning

engineers are at the center of that

transformation. You might be thinking

okay I get it machine learning is

impactful but what exactly does a

machine learning engineer do compared to

other roles in tech? That's a great

question. In the world of machine

learning, you'll hear about a few key

roles such as data scientist, machine

learning, and AI engineer. Let's break

them down so you know exactly where you

fit in. Data scientists are like the

detectives of data. They spend their

time analyzing large data sets, finding

trends, and trying to extract meaningful

insights. They build models, but their

main focus is usually on data

exploration, and experimenting with

various algorithms. They don't typically

focus on deploying those models into

production environments. Machine

learning engineers on the other hand

these are architects. They take the

models built by data scientists and

build scalable deployable systems. They

work on creating solutions that will not

only work in the short term but can also

scale to handle real world data in

massive volumes. The machine learning

engineer is responsible for ensuring

that machine learning systems are

integrated into businesses that can work

seamlessly with existing technologies.

AI engineers focus more on the

application side of things. They build

AI powered products like chatbots, voice

assistants, and real-time systems. While

their work often overlaps with ML

engineers, they are typically more

focused on the userfacing product and

how machine learning fits into it. As an

ML engineer, your primary focus is to

take models and turn them into

actionable solutions that are deployed

in real world systems. We shall now move

on to why 2026 is the right time to

enter machine learning. Now that you

know the role of an ML engineer, let's

talk about why 2026 is the perfect time

for you to jump into the field. You've

probably heard that machine learning is

a hot topic, but what does that mean for

you as someone starting out in this

field? So, I'll help you break that down

for you. First, the demand for ML

engineers has skyrocketed. The world is

full of problems that needs solving, and

machine learning has proven to be one of

the most effective tools to solve them.

From predicting customer preferences to

automating critical business functions,

machine learning is changing how

businesses operate. Secondly, the tools

used to build machine learning systems

are more accessible than ever before. In

the past, machine learning was viewed as

something experimental, something that

required a lot of effort just to set up.

But today machine learning platforms and

frameworks such as TensorFlow, PyTorch

and Scikitlearn have matured

significantly. These tools make it

easier to build and deploy models that

can scale to handle real world data. We

shall now move on to why is this the

right time for you to get started out as

a machine learning engineer. So how do

you get started? The first step is to

build a strong foundation. You might be

excited to start building models and

diving into algorithms. But before that

you need to understand the core concepts

that drive all the machine learning

systems. These include mathematics,

programming and data handling. You don't

need to be an expert in all of these

areas, but you do need to understand the

basics. Think of these as building

blocks of everything that you will need

to learn machine learning. Let's start

with mathematics. You don't need to be a

math genius, but you do need to

understand the basics. There are three

main areas of math that will help you

get started out as an ML engineer, and

those are linear algebra. This is the

study of vectors and matrices which are

used to manipulate and process data in

machine learning models. If you have

heard terms like feature vectors or

matrix operations, that's linear algebra

play. This is the core of how machine

learning algorithms operate. This is the

math behind optimizing models. You'll

use calculus to adjust the parameters of

machine learning models and minimize the

error between predicted and the actual

values. Specifically, derivatives are

used to find the best way to fit a model

into the data. Statistics. Machine

learning is all about working with

uncertaintity, and statistics will help

you make sense of it. Whether you're

dealing with probability distributions

or hypothesis testing, statistics can

help you understand patterns in data and

make decisions with uncertain

information. These mathematical concepts

will help you build a more accurate

model and optimize it effectively. Let's

move on to programming stack. If you're

new to programming, don't worry. Python

is the language that you will want to

learn. It's simple to get started with

and has a huge ecosystem of libraries

especially designed for machine

learning. The key libraries that you

will need to master include NumPy. This

library is used to handle large arrays

of matrices and data which is the

backbone of most machine learning

algorithms. If you plan on working with

large data sets, you'll be using NumPy a

lot. Pandas. This library is great for

data manipulation and analysis. You'll

use pandas to clean, organize, and

transform data, making it ready for

machine learning models. Scikitlearn.

This library provides simple, easy to

use tools for building machine learning

models. It covers everything from data

prep-processing models like regression

and classification. Once you're

comfortable with these tools, you'll be

able to start building machine learning

models and working with real world data.

Next up is SQL, which stands for

structured query language. As an ML

engineer, you will be working on lots of

data, and SQL will help you query

databases to retrieve information that

you need. You'll use SQL to extract

data, filter it, and join tables

together, making sure that you have the

right data to your models. Once you have

your data, the next step is data

wrangling. Data wrangling is a huge part

of your job, and if you master it, it

will save you a lot of time and

frustration while building models. We

shall now move on to types of machine

learning. Let's talk about the different

types of machine learning. There are

three main categories which are

supervised learning, unsupervised

learning and reinforcement learning.

Speaking of supervised learning, this is

when you have label data. You train your

model on data where the answers are

already known. The goal is for the model

to learn the relationship between inputs

and outputs so it can predict the future

outcomes. Unsupervised learning. In this

case, the model works with unlabelled

data and it tries to find hidden

patterns and groupings in the data. This

is useful for tasks like clustering or

anomaly detection. Reinforcement

learning. This type of learning involves

training an agent to make decisions by

interacting with its environment and

receiving rewards or penalties based on

its actions. This is often used in game

AI or robotics. We shall now move on to

feature engineering and evaluation.

After cleaning your data, the next step

is feature engineering. This is the

process of transforming raw data into

meaningful features that help your model

make better predictions. Once you've

engineered your features, it's time to

evaluate the model. You'll use metrics

like accuracy, precision, and recall to

see how well your model is performing.

Proper evaluation ensures that your

model is ready for real world

applications and that it can handle data

in production. We shall now move on to

the six-month learning plan in order to

become an ML engineer in 2026. So, how

do you actually get started? Here's your

six-month learning plan to guide your

journey. Firstly, focus on learning

Python, mathematics, and SQL. Dive into

machine learning algorithms and hands-on

projects. Then, learn deep learning and

work with frameworks like TensorFlow or

PyTorch. By the sixth month, you can

work on a capstone project that covers

the full pipeline from data collection

to deployment. We shall now speak about

the ML engineer tool stack. So let's

talk about the tools that you will need

as a machine learning engineer. You must

be wondering what tools should I be

learning to become a successful ML

engineer in 2026. So let's simplify

this. Git and GitHub are the first

things on your list. At first glance,

version control might seem like

something that only coders need to worry

about. However, it can be really

crucial. Why? Because version control

allows you to track every change that

you make to your code, collaborate with

teams, and roll back changes if

something goes wrong. Imagine you're

working on a huge project and you mess

something up. Without version control,

you could lose hours of work. But with

Git and GitHub, you could go back and

fix to any point and time. You might be

thinking, okay, I get that Git is

useful, but what about the tools that

can actually help me build and deploy

machine learning models? Now, that's a

great question. So let's talk about the

cloud platforms like AWS, Google Cloud,

and Azure. In the past, machine learning

models were often built and tested upon

local machines, but we quickly realized

that it wasn't scalable. These cloud

platforms give you the ability to handle

large data sets, run models on

highowered servers, and scale them as

your projects grow. Instead of relying

on your personal laptop to do all of the

heavy lifting, you can leverage the

cloud to train models much faster and

handle massive data sets without

worrying about memory limitations. We

shall now move on to the next topic

which is on MLOps life cycle. You have

now got all of your tools in place. But

how do you go from all of these tools

coming together in the MLOps life cycle.

So you must be wondering what does all

of the MLOps life cycle tools even look

like and how do I manage the process

from starting to finish? Well, here's

the thing. MLOps is like a welloiled

machine. It involves a series of stages

that ensure that your models remain

reliable, scalable, and adaptable to new

data over time. Think of it like a car

assembly line. First, you train the car

and then you deploy it into the real

world. After that, you need to monitor

it and see if it runs smoothly. If it

breaks down, you will have to retrain

it. Let's break it down into key steps.

Training your model. This is where you

create the model using your training

data. This is a very experimental phase

where you try different algorithms,

tweak parameters, and optimize the model

for better performance. Deploying it to

production. Now that your model is

ready, it's time to put it into action.

This means making the model available to

users or clients. Whether it's in a

mobile app or a web service, deployment

is where the magic happens. Monitoring

performance. Once your model is live,

you can't just forget about it. It's

like checking the car tire pressure

after it's been on the road for a while.

You need to continuously track how well

your model is doing. If it starts to

slip or underperform, it's time for

tweaks and adjustments. Lastly, we have

retraining. Over time, your model may

need to be retrained with new data to

stay accurate. This is especially true

in industries where the environment is

constantly changing like e-commerce or

finance. Restraining ensures that the

model stays relevant and continues to

provide value. You may be wondering that

this sounds like a lot of work and

that's true. That's where MLOps tools

like MLS help you streamline this

process. By automating and managing this

life cycle, you can spend less time

dealing with the back end and more time

focusing on building innovative

solutions. We shall now move on to

experiment tracking. As you dive deeper

into machine learning, you'll quickly

realize how important it is to track

your experiments. At first, this might

seem a little overwhelming. You might be

thinking, I can't just run a model and

hope for the best, right? But trust me,

tracking your experiments is one of the

best habits that you can develop early

on. Think of it like logging your

workout progress. Even if you don't

track the results, how will you know

whether you're improving or not? The

same goes for machine learning. By using

experiment tracking tools flow and

weights and biases, you can keep a

detailed record of every model and every

hyperparameter and every evaluation

metric. Now imagine you're building a

recommendation system for an online

store. You try different adjustments,

algorithms, and get different results.

With experiment tracking, you can easily

compare which configurations work best.

You can easily compare which

configurations work best and learn from

the past mistakes. You'll always know

which experiment gives you the best

results and which needs tweaking.

Tracking experiments also helps in

collaboration. If you're working with

the team, being able to see everyone's

experiments in one place makes it easier

to understand their approach and build

on each other's work. We shall now move

on to projects and portfolio. Now that

you've got your tools and workflow in

place, it's time to focus on building a

strong portfolio. You might be thinking,

how do I make my portfolio stand out to

potential employees? That's a great

question. The answer is simple. Real

world projects. Think about it.

Employers want to see what you can do in

practice. They don't just want to see

theoretical knowledge. They want to see

that you can solve problems and build

working systems that can make a

difference. So, for that reason, we have

mini project examples. At this point,

you're probably itching to start

building something yourself. Well,

hands-on projects are the best way to

solidify what you've already learned.

For example, you could create a customer

churn prediction model or a fraud

detection system. These are practical

real world projects that demonstrate

your ability to build solutions from

start to finish. Make sure you document

your projects clearly on GitHub and

always include a detailed explanation of

your approach, challenges, and results.

Remember, the goal is to show that you

understand the problem and build a

working solution. We'll now move on to

Kaggle and open source. As you continue

to build your portfolio, I highly

recommend diving into Kaggle

competitions and contributing to open

source projects. Kaggle is an amazing

platform where you can work on real

world data sets and solve problems that

companies and research institutions are

facing. Not only will you improve your

skills, but you'll also have the

opportunity to see how top data

scientists and machine learning

engineers approach similar problems.

Contributing to open-source projects is

another excellent way to showcase your

skills. It shows that you can work well

with others understanding existing

systems and contribute to the community.

Plus, it's a great way to gain

visibility and make connections with

other engineers. We shall now speak

about the rรฉsumรฉs which work for you.

Speaking about your resume, when it

comes to landing a job as a machine

learning engineer, your resume needs to

be focused on real world projects and

practical experience. So, be sure to

highlight the projects that you've

worked on, the tools you've used, and

most importantly, the impact that your

work has had. If you build a

recommendation engine that boosted

product sales by 20%, make sure that you

include that. Employers want to see how

your work contributes to solving real

business problems. Don't forget to

include links to Kaggle profile or any

other open source contributions that you

have made. Employers love seeing code

and showing them your projects is the

best way to stand out. We shall now

speak about what interviewers test in

2026. By now you must be wondering what

do employers actually look for in an ML

engineer. So in 2026 interviews are not

just about technical knowledge.

Employers just want to see how well you

can communicate your thought process and

real world problems. You'll likely face

practical tests that challenge you to

build a model or analyze data in real

time. You'll also be tested on how well

you explain your approach and justify

the decisions that you made during the

project. Prepare to discuss things like

why you choose a specific algorithm for

a problem or how you handle issues like

data imbalance or overfitting along with

what metrics you use to evaluate your

model's performance. Being able to

communicate clearly your process and how

you arrived at your solution is a skill

that will set you apart from other

candidates. We shall now speak about the

ML trends that you will need to follow.

Machine learning is evolving fast and

staying up to date is key to remaining

relevant in this field. So here are a

few trends to keep an eye on in 2026.

Generative models. These models can

generate new data based on patterns they

learn and they're used for things like

text generation, image creation, and

even music composition. AutoML automated

machine learning tools are making it

easier for non-experts to build machine

learning models. As a result, more

people will be able to contribute to

this field without needing to become

experts. Privacy first models. As

privacy concerns grow, machine learning

models are being designed to work

securely and ethically and these don't

compromise on user privacy. Staying on

top of these trends will help you remain

competitive and innovative in the ever

evolving field of machine learning. So

in conclusion, I will say consistency is

key in machine learning. You don't need

to know everything right away, but you

must stay committed to learning and

building. Start small, keep

experimenting, and keep improving. The

road to become a machine learning

engineer may seem long, but with the

right tools and mindset, you can get

there. Thank you for watching and I'll

see you in the next one. Thanks for

watching the machine learning engineer

road map for 2026. We hope that this

video gave you a clear path to becoming

a successful ML engineer. Ready to take

the next step? Explore the Simply Lance

professional certificate in AI and

machine learning in partnership with

Purdue University. Gain hands-on

experience and industry recognized

credentials to boost your career.

>> Machine learning is which is a subset of

artificial intelligence, right? That's

uh basically um machines learning from

data

in order to uh make decisions

essentially. Um, so this was a big

departure from the rules-based systems

at the time, right, that were explicitly

programmed to make decisions. So just

think of an example like a really big

kind of if this, then that, then that,

then that, and and else if this, this,

this, right? So bunch of rules that had

to be pre-programmed in order to um come

out with some final answer. Uh, with

machine learning, it's the exact

opposite of that. we're actually

training something from examples from

existing data um in order to predict

something or um

make some type of decision. Uh and so

we're going to learn about the various

ways we can do machine learning. But if

you guys remember we

um talked about some of this like the

differences and the uh basically

rules-based approaches to learning from

data approach. Um and in included in

that is going to be uh complex

unstructured data. So things like

images, text, audio. What handles those

really well is uh deep learning which we

will get to in the course after this.

But uh those are certainly in there as

learning from data even complex data.

So we had this picture uh and I think

this is kind of around where we left off

last time was uh just distinguishing

between those three terms. We see

artificial intelligence, deep learning

and machine learning kind of used

interchangeably, but this is really how

they fit in. Artificial intelligence is

kind of a broad anything mimicking human

intelligence. Um which doesn't have to

be learning from data, but uh machine

learning is part of that. And then um

one way to accomplish machine learning

is to use neural nets which is the focus

of uh deep learning. Um and so deep

learning has been has found a lot of

success especially recently with uh

those complex data types like images,

speech, text, right? So deep learning

used all over the place. Even in um

modern like generative AI, we see deep

learning used quite a bit. Um it really

anything that's using neural nets is uh

going to be deep learning.

Um

and again we'll focus on that later but

we're going to be mainly focused on

machine learning for this course

primarily machine learning that does not

use neural networks. Okay so just models

that are not necessarily neural networks

be our focus.

So in machine learning we had an example

of a game uh essentially um learning

what decisions to make uh based on the

uh kind of current um state of the

board. This could be a um you know

machine learning example that uh learns

from many previous examples. So a lot of

data around these games are used to

train these um kind of robots that can

play these games and play them at a very

high level. Um so there's been a lot of

successes actually in machine learning

and deep learning um around

uh playing games like chess or go

um using machine learning algorithms. So

pretty cool.

All right. So I think this is where we

ended. We last time we said there's a

bunch of different use cases for machine

learning. So um recommendation system is

going to be a big one and we will

actually study that uh in one of our

final lessons of this course. Um chat

bots like generative AI doing sentiment

analysis chat bots we'll study later but

those are certainly an application of

learning from data in order to uh

generate responses to text prompts

right. Um spam filtering that's a good

example like classifying an email as

spam or not spam. Um that that gets

trained from examples and uh learning

from data such as previous emails. Um

social media posts analysis is another

kind of text data um use case but you uh

can do a lot with that text like you can

predict the sentiment um you can predict

uh the category of what what the post is

talking about um those kind of things

all can be done with machine learning

>> and many other use cases not on this

list that we will uh cover

>> you know as we as as we go further.

Okay, so this is where we kind of left

off. Um, so what's doing all the hard

work here is

>> uh machine learning algorithms. So these

are things that will um these are things

that will learn from the data. So they

are uh they they are basically um

algorithms or sets of rules that uh or

mathematical rules I should say not

formal rules like in the in the sense of

a rule system but mathematical um

formulas and mathematical uh rules

essentially that help us learn from the

data. So they correlate the data to some

type of outcome. So some type of

prediction uh whether that's going to be

as we will see whether that could be

like a number like we're predicting a

price or demand or sales

um or it could be a category like is

this transaction fraud or not fraud or

what's the probability that this is

fraud um so we have different kinds of

predictions we can make with machine

learning.

Um

but uh we will study the kind of the

differences of those coming up. Um but

machine learning algorithms are really

what power they're kind of the models,

right? They're the models that help

power uh machine learning to actually

learn from data.

So we're going to spend a lot of time in

this course studying those algorithms

like the different models that we can

build and what their differences are,

what their strengths are, what their

weaknesses are. we'll we'll learn a lot

about those.

Okay. So, I guess you can imagine like

everything is so data dependent, right?

Um we're learning from data. So, uh it

makes sense that the quality of data

really really matters here in

determining how strong the model can be.

Um so you see this graph here charting

kind of the um high quality data um

versus just uh any old data but a decent

enough quantity of it. Um you can see

that performance and the performance is

measured by some evaluation metric. Um,

so think of it as uh something like an

accuracy. Like if we were predicting

fraud or not fraud, how accurate can our

model get at actually detecting fraud

um it gets better and better and better

the graph shows that the higher quality

of data that we have. So there's kind of

that there's a there's a saying in

machine learning um called garbage in

garbage out. What that means is if you

have poor data, even the best model in

the world, poor data is not going to

result in having a good model that can

be accurate and perform well. Um, so it

needs to be high quality, meaning um

there needs to be a decent amount of it

and it needs to be labeled appropriately

as we will will talk about

um and it needs to not have any, you

know, significant outliers. it needs to

be clean, not have those missing values,

all of those things. Um, you can you

have a good chance at deriving good

predictions from higher quality data

as this kind of shows.

Okay.

So, one thing we're going to learn um as

we go along is

quantity matters as well. So, not only

quality, but a decent amount of it. And

um we're going to learn those kind of

rules of thumb like how much data do I

need for certain algorithms. Um one

thing that we will see is that uh the

the basic machine learning models that

we'll study don't need as much as a

neural network would. It you know neural

networks are going to require a lot more

um than a basic machine learning model

learning model. So uh that's something

we will see as we go along. But uh this

is something we'll talk about and

discuss with each model that we study is

kind of how much data do we actually

need to produce a high quality model.

Okay, any questions uh so far?

Okay, let's talk about the different

types of machine learning that we're

going to discuss. Prim, there's going to

be two primary ones that we will study

in this course and then a couple others

that'll be a little bit more advanced

that we won't get to but worth knowing

about. Um, so there's going to be four

total that we'll study or talk about and

they'll be on this list here, which is

um supervised learning and unsupervised

learning. Now, I'd say the majority of

our focus will probably be on supervised

learning, and we'll talk about what that

means, but we'll also cover unsupervised

learning as well. And so, we'll look at

the most popular techniques in each of

these types of machine learning.

Um,

and then we'll talk about these two, but

not really study them because they're

more advanced topics. Um, that that will

be beyond the scope of what we'll do.

But uh these are going to be um

different styles of machine learning

that are going to be characterized by um

what kinds of predictions they make,

what kind of data they need and require.

Um and uh what kind of outcomes they're

actually producing. Um, so let's let's

get into each of these, but uh the the

one that we'll probably spend the

majority of our time on is going to be

supervised learning, but we will study

unsupervised learning as well. We'll

study both and we're going to talk about

we're going to define both of those um

coming up. And again, these will be a

little bit more advanced topics that we

won't spend too much time on.

Um, but but we'll discuss their

relevancy in machine learning. Um, and

give a good definition to it.

Okay.

All right. Let's start with supervised

learning. Now this is going to be uh a

term that really refers to

using examples. So using labeled

examples. So here we say labeled data to

help our model train. In other words,

help our model be able to predict guided

by specific input output pairs. So

supervised really refers to the fact

that we have answers. We have examples.

We have answers with those and we use

that collection of data to build our

model off of so that we can predict

um those kinds of things like a price,

like a category, like a spam, not spam.

in this in this slide like we would be

predicting if this shape is a square, a

triangle or a circle.

Um but but when we build a model for

that, we have data that has an answer

attached to it, right? We've talked

about this before a little bit with

labels. So there's a guide there that

can guide us towards building our model.

there's an actual every every example

has an answer and that answer is really

critical to help build our model off of.

So, um that's it's almost like you have

um a you have a bunch of exercises

in let's say like a math textbook. You

have a bunch of exercises and you have

the answers and that way you can kind of

check your work. you think about model

training um that is the really a lot of

that process of model training as we are

going to discover is um basically

checking our work against these answers

in our data in our training data.

Okay. So supervised learning is any type

of machine learning that involves

learning from labeled data in order to

predict outcomes. Okay. Predict outcomes

like now the the outcomes can be

numerical. They can be like a price,

temperature, demand, sales, revenue.

They can be numerical, but they can also

be categorical. So they can be like spam

not spam, fraud, not fraud, cancer, not

cancer. Um dog, cat, giraffe, those kind

of categories. Um we could predict

those. It's some type of outcome. Okay,

some type of outcome. The key is we're

using labeled examples to guide our

model building. That's why it's called

supervised learning.

So we know in our data we know what the

inputs are. Of course, those are going

to be think of the inputs as like all of

our columns and then we have a special

label column that represents the output

we're trying to predict. So if you think

about that housing price data, the label

could be the price. And that's something

we would build a model to predict, but

we have answers for all of our examples

in our rows. We have answers to help

guide our model building.

They help tweak our model because we

know the answer ahead of time. So

they're they're really good examples to

build our model off of.

Okay. So that's that's supervised

learning.

Uh in this example, is circle not in the

prediction because it's not part of the

test data even though it's in the

labeled data?

Um no, it just not necessarily. It just

means that like we learn against all of

these examples that have these answers

and then when we observe new examples um

we can try to predict what those would

be based on what we've seen before. So I

if there was a you know it's just a

coincidence we only have two two

examples in our test data like we could

have a circle here in which case we

would predict circle

that's fine or at least we would hope

our model would predict circle right

that's what we're hoping may or may not

get it right

um but it's it's only not there because

we only like we're just assuming that we

only have two examples we're testing

against but in reality we would probably

do a lot more than two.

It's just it's just a coincidence

really.

In reality, we would test against a lot

more data. And we're actually going to

see why we would do that. Like why would

we train our model and then kind of use

additional data um to to evaluate it?

It's actually really important that we

do that step to get a sense of how good

our model is before we take it out in

the real world. So if we apply our model

that we build on our label data to

um this kind of set of test data that we

haven't been exposed to before. It helps

give us a sense of how good is our

model. So it's test is usually used for

evaluation.

So that's something that's something

we'll study.

How do we train? Uh it depends on the

model. Um so training will be a sense uh

will be an algorithm that will um

basically update the model according to

the data. These labeled examples. Um

every model is going to be different in

exactly how it trains. So we're going to

we're going to talk about that when we

get to the individual models that we'll

study.

But uh loosely speaking, they're going

to use the data to adjust itself. Like

imagine adjust like tuning a bunch of

knobs. Um, like the best example I can

give you is we I think I did this one

last week where you have kind of a

function

that predicts the price and let's say it

has

um weights like weight one with feature

one, weight two with feature two,

weight three with feature three. So

imagine we had three input features and

we we built an answer according to that.

Essentially what we would do to train

the model is adjust these

um in order to get this correct based on

our our labeled examples.

Okay.

So that's something we're going to learn

about coming up shortly when we when we

actually dive into model. Every model is

going to be slightly different in how it

trains, but at a high level it's going

to use the training data with those

examples, right? the labeled examples to

help guide the formula essentially to

adjust to generate the proper kind of

model here. The these things are going

to be adjusted according to the data

in order to produce the correct output.

So think about these as knobs that will

turn.

Okay.

Uh which type of machine learning is

used? Uh probably supervised um which is

what we're talking about now. So

probably supervised because most people

want to

um build some type of model to predict

something. Uh so yeah, I'd say I'd say

supervise.

Yes, we're are we are definitely going

to learn how to train. Yeah, we'll see.

We'll do the code. Um, I'll tell you

about how it's done. Yeah, we're

definitely going to learn it. But what I

was saying is it's kind of on a model by

model basis.

So, I want to wait till we get into the

individual models, then we'll talk about

how they're trained.

But yeah, we'll we'll learn how to do

that.

But yeah, supervisor is used all over

the place. Even even for uh generative

models, they use supervised learning

because um like an LLM

is going to use labeled examples in

order to train, right? In order to train

how to generate responses according to

prompts. Um it needs to learn against a

lot of text examples.

So that supervised learning is what um

results in that model,

right? Learning from those labeled

examples.

Okay.

It is yeah image image uh a lot of um

yeah a lot of image processing is

supervised like object detection. So the

yolo model is an object detection model.

Yes. Um because it has to be trained

right it has to be trained on uh it has

to be trained on images

with labels such as this is what object

is in this image. This is the box around

the object.

Um yes. So if if it's if it ever uses

label data to train and build the model,

it is supervised. So YOLO is definitely

supervised and we actually we will we

will cover the YOLO model later on in in

our deep learning course. We talk about

object detection.

So we'll we'll study that.

But yeah, it's supervised

Okay. So on the slide we have some

common supervised learning algorithms

that are we will study. So all of these

we will study and understand what they

do and how they work but just giving you

some to name them. linear regression is

kind of the one I just drew out which is

the um this is the prototypical like

easiest to understand model that is kind

of the um exactly like this where we

have a weight times a feature um a

weight times a feature and then a weight

times a feature

and on and on and on. You can have as

many as you want.

um that is a linear regression. And so

that is um that's a supervised model

because we need this value here and we

need all of our inputs in order to um

actually train this model and generate

all those weights

um that that is uh that uses um labeled

examples to help tune all those knobs.

Um same with all these other models. So,

we're going to talk about decision

trees. We're going to talk about

logistic regression and and SVMs, which

are support vector machines. We'll talk

about all of those, but they're all

examples of supervised uh supervised

learning.

Okay, we'll talk about all of these.

They're all supervised because they all

require labeled examples in order to

train them and and then subsequently use

them. Okay.

Okay. So what are some use case

examples? So for for instance in uh

supervised learning we may be predicting

temperature based on yearly temperature

trends. So we would have that yearly

data as our um as our labeled examples

and those would supervise the learning

of a model that predicts temperature.

Um, same thing with predicting crop

yield based on um, seasonal crop quality

changes. So maybe we have a bunch of

features relating to crop quality. We

could predict crop yield. Um, we would

just need historical examples with those

labels, right? What the crop yield is

for each time period. Let's say we would

just need those uh, supervised examples

and we could easily build a model off of

it.

Um

uh this this last one sorting waste

based on known waste items and their

corresponding waste types. Um that's

kind of like spam. It's like filtering

basically like a spam filtering. Um so

think of it like the the shapes example.

We sorting things into squares, circles,

triangles. Um, same kind of idea here

where we have a bunch of examples on

what those um what those waste items

should uh should belong to, like what

waste bins they would go to, for

example. Um, and those could be labeled

and therefore then we could um

understand what category of waste they

belong to.

Um, same thing with spam. something is

fraud or not fraud, spam or not spam,

cancer or not cancer. All of those are

going to be supervised learning examples

because they're going to require in

order to train them, they're going to

require data that has those labels.

Okay? So, anything that has labels is

going to be supervised learning.

So, again, this is where we will spend

probably the the majority of our time is

doing supervised learning problems. ones

that we have labeled data. We're

building a model and we're going to

predict those those uh labels

essentially.

Okay, before we go to unsupervised, any

questions about uh supervised

Okay.

All right. So, supervised requires

labels

in order to have an example to go off of

to build your model. And that's because

you're predicting those kind of outcomes

like spam or not spam, cancer not

cancer. Now unsupervised learning is

completely different. It's the opposite.

So unsupervised learning is where we do

not use labels whatsoever. So we're not

using any labels at all. So it's it it

can be completely unlabeled or even if

it's labeled, we're not using labels in

any way. But um we primarily would say

it's unlabeled data. We have no guidance

because we're not using the labels in

any way. we have no guidance to um

predict anything but that's because

we're not really predicting anything in

unsupervised learning. Generally what

we're doing is looking for some

structure or pattern.

Okay, with unsupervised learning we're

looking for some structure or pattern.

So um one type of example that's very

very popular is going to be this second

one which is um identification

identification of user groups based on

similarities or commonalities. Now this

is going to be a problem basically known

as clustering

and it's a problem we will study quite a

bit. There's going to turn out to be

lots of different algorithms that can

accomplish clustering. So what

clustering attempts to do is basically

say um we have data that's like this and

then data over here and then data over

here. Let's just group these together.

So like this should be one group. This

should be one group and this should be

one group. And we can find those

structures and say okay this is group

one this is group two and this is group

three.

One 2 3. And we can basically build what

we would call clusters of data um based

on how close together the points are

kind of located in these kind of cluster

zones like these boxes I've drawn.

Okay. Now that doesn't require any label

to do which is really fascinating. So

unsupervised you don't need any label at

all to accomplish the algorithm. Um so

clustering is one good example. Um

finding outliers or anomalies is

another. So we don't necessarily have

any label of what is an outlier or what

is an anomaly. We are deriving that from

the features alone. There's no guidance.

There's no label um to doing like

outlier detection or anomaly detection.

Okay. So that's another good example.

One that's not listed on here um but is

also really important that we will study

is something known as dimensionality

reduction.

So dim reduction and what that what this

focuses on is basically compressing the

data set a bit. So we take our data and

basically compress it. Um so that but we

do it in such a way that we retain as

much information as we can. It's a very

like smart compression and what it does

is it lowers the dimension.

um dimension. Think of the dimension as

like number of columns.

Number of columns.

So imagine we had 100 columns in a data

frame. What we could do is actually

reduce that down to 10. So like 10% of

that. So we reduce it down to 10. And um

but those 10 are it's not like we

chopped out um 90 other columns. we um

smartly kind of compressed all that

information into these 10 new columns um

that are compressed versions of the

hundred that we used to have. Um so

dimensionality reduction is is another

unsupervised technique. It requires no

guidance, no label to do, but is um a

really useful technique to reduce the

size of your data if you're doing things

with it. Um so this is another one that

we will we'll study how to do it and

basically more details behind it what

the algorithms are.

Um we'll so probably those two in

unsupervised will spend the most amount

of time on clustering and dimensionality

reduction.

Uh and supervised if some data is

present but we didn't label it means in

example we had circle triangle square in

the training data we add pentagon

but we didn't label that in that case.

Uh yeah. So every um in supervised

learning, every row, think about it as

like every row in our data frame needs

to have a label

uh associated to it. It needs to have a

a column that represents the label.

So if we've never seen Pentagon before,

I can't use that as a label.

So it has to the pentagon has to exist

in the data. if I'm going to be able to

predict it,

right? So, it can't predict, right? If

we've never seen it before, we have no

examples to go off. We have no guidance.

So, how could we predict that?

Right? We can't predict it

if it's if it's in there. So if if we

have labels of Pentagon, let's say, then

yeah, we could predict Pentagon.

We could

remove. Remove what?

We wouldn't if it was talking about the

Pentagon, we wouldn't remove that. No,

let me go back to that page. We wouldn't

remove it. Um, it's just if it's not in

our labels, we're not going to be able

to predict it. So, Pentagon's a good

example here. Uh, Pentagon is not one of

our labels. So, it currently is not in

our data set as one of the labels. We

only have data that's either a triangle,

circle, or

square. We don't have pentagon. So, I

would never be able to predict pentagon.

I'll never be able to do that if I

haven't seen examples of it before.

Okay. But let's say we had that in

there.

So, we had Pentagon.

So if we had Pentagon, um we could have

an example of it in our labels

and then Yeah, we it could be then we

could predict it.

Yeah. Yeah. The the don't get worried

don't worry about the test data. So the

test data is just saying here's a new

here's a shape what is it okay that's a

square here's a shape what is it okay

that's a triangle and we could have as

many of those examples as we want in our

test data so we could have a circle and

say okay what's this should be circle

right the test data can be whatever it

whatever it wants but yeah if if we've

never seen pentagon before we're never

going to be able to predict

These are the the label data and labels

are basically the talking about the same

thing. The labels just mean what are the

categories

that are present in our data. So in this

data we only have three labels that are

present.

So the labels is are relative to our

label data, right? It's saying

what labels,

excuse me, what labels uh do we have

in our data and we only have those three

circle, triangle, square. So so Pentagon

would not be part of those labels. We

couldn't predict it.

No. So unsupervised is not going to make

a prediction. That's the big difference

with unsupervised. They're not going to

make a prediction like this. Um so

unsupervised is not going to make a

prediction. It's going to do something

different like um basically say like

these guys are similar, these are

similar, these are similar, this is a

cluster, this is a cluster, this is a

cluster. It's not going to make a

prediction. That's what supervised

learning does.

Clustering, yes, which is unsupervised.

Yes, clustering does not require any

labels. Unsupervised just means we don't

have any labels. We don't require any

labels.

So the other thing unsupervised might do

is it might say

and again without the labels it might

say that this is an outlier.

it might say that this guy is an outlier

because there's only there's only one of

those and they're not like the other. So

that that's something that um that's

something that uh unsupervised could do.

Um it it yeah and no. It kind of labels

a cluster in the sense that um it would

basically assign a number to it like

this is cluster one, this is cluster

two, this is cluster three.

It'll assign a number to it, but it's

not a very meaningful it doesn't assign

like a prediction label in in the

traditional sense of a label.

It does provide like a numerical index

for the cluster to because what we want

to know is like okay this guy has the

cluster of one. This guy belongs to

cluster one. This guy belongs to cluster

one. This guy belongs to cluster two.

This guy belongs to cluster two. Does

that make sense? So there needs to be

some like index of what cluster you

belong to.

So it's kind of like a label but not in

the traditional like prediction sense.

Okay.

Very good. So again, unsupervised, no

labels. You're doing things like

identifying clusters,

um identifying outliers, doing

dimensionality reduction. These are all

like structure and pattern oriented

things. They're not predictions of a

label. Okay? They're not which is what

we would see in supervised learning.

Okay. So an example would be that we

take we put in the data um we can group

together uh data such as images into

categories based on similarities um

which would be like those clusters. So

there's no these would be groups that we

don't have any label on ahead of time

like we don't have we don't say that

this image should belong to this this

image should belong to this we derive

that from the characteristics of the

data. Um so think like a good example is

um customer groups. So we would identify

customers based on like okay do they

have similar spending levels? How many

days do they go shopping in a week? How

much money do they spend? And we can

kind of group together customers based

on similar qualities.

Clustering will find those groups that

should exist.

um it will discover those groups based

on um the similarities in the data, but

there's no labels that that say like

this person should be in this group,

this person should be in this ahead of

time. There's no labels of that. It gets

derived during the algorithm. It's

unsupervised,

right? There's no unsupervised really

literally means no guidance. There's no

guidance to doing it. We just derive

that from the structure of the data

which is the similarities.

Okay.

All right. So,

a couple more for you. So we had um

supervised which uses the labels. We

have unsupervised which uses no labels

looking for structure. And then we have

something that's kind of in between

which is um what is known as

semiupervised learning. And this is

where you use a combination of a little

bit of label data, but most of your data

is actually unlabeled data. Um, and you

try to get some use out of that label

data in order to um build a model out of

it. And so, uh, it uses the, um, it uses

that label data to, um, generally

provide some guidance on usually what

happens with semi-supervised learning is

you use your label data to kind of

predict what the label should be for the

unlabelled data and then you can go from

there. So you can create artificial

labels on this unlabeled data and then

you can use all of it once it's all been

labeled kind of like a supervised

learning uh approach. So but but this is

semi-supervised basically refers to the

fact that you start out with most of

your data not being labeled but you do

have some labeled examples and what you

can do is basically extrapolate those

labels into the unlabeled data set and

then provide some artificial labels and

then now everything has a label you can

do supervised learning.

Okay. So, it falls kind of between um

supervised and and unsupervised.

Uh and there So, this is this is kind of

rare. Most of the time you're not going

to do that. You're actually just going

to um prefer to just start with all

label data. That's usually the preferred

approach. Most of the time you'll

actually just be doing supervised

learning, not really semi-supervised

learning. So, it's pretty rare, but um

it it could like if Yeah, it could if

the if we had a lot of examples of

Pentagon and we wanted and so they were

unlabeled and then we tried to guess

what kind of shape they were um and

provide an artificial label uh and then

um then use that whole data set to build

a model off of then then yeah, it could

it could fall into this category. Okay.

They Oh, going back to the question,

they still use some kind of label data

like age, gender. They use uh that's

those aren't those aren't really labels.

That's the features. So, yeah, they

still use the core features of the data.

They just don't have any like labels in

the traditional sense of a label. Like

you should think of a label as something

we are trying to predict.

So whether that's a price, whether

that's like a category like spam, not

spam, cancer, not cancer, it's something

we'd be interested in kind of

predicting. And so um in our data, we

would have an answer for every row. We'd

have one of our columns would be like

the the result like the outcome answer

that we're trying to predict. That's the

label.

So in unsupervised, we don't have any of

the labels.

We do have just the regular features

like gender, age, income, square

footage,

bedrooms, bathrooms, all those things.

Okay.

So, we have semi-supervised that falls

in between supervised. Now, the reason

it falls between is be is because

there's a decent amount of data that's

unlabeled. In fact, a majority of it

unlabeled. But what we can do is try to

label it. We can try to take what we

know from our existing labels and

predict an artificial label and then use

all that data together in kind of a

supervised fashion for a model down the

road.

So that's kind of what this picture uh

says is we can try to take um you know

maybe we try to infer some labels based

on we have some some labelled data here.

We have most of our data is unlabeled

and we try to supply some labels to it.

Um like maybe we have a baby's category

of teens, a tween, uh you know youth and

um adults. Um and then we try so we we

take our our labels and we try to

extrapolate those into artificial labels

for this unlabelled data so that we can

use it now because then everything has a

label at this point and then we can just

go ahead and do supervised learning from

there.

So we can do supervised from there. What

we would prefer to do and what we'll do

in this course

um is just start with supervised. We'll

just start with the labels. We won't try

to derive artificial labels usually.

We'll just start with labels.

So one example in the real world is

something like Google photos which um

whenever you take a picture it can

provide uh uh labels based on previous

uh images in your library. So it can it

can produce tags or um labels on those.

Uh generally when you take that picture

it's kind of unlabeled unless you go in

and specifically provide some tags and

some labels. But um if you don't do that

it can still it can still uh make it can

artificially create one of those based

on the other label data that you already

have.

So that's um

that's an example.

Okay.

All right. Last one in terms of machine

learning. So we have supervised, we have

unsupervised.

Uh then we had semi-supervised which is

somewhere in between a mixture of having

some unlabelled data and label data. Um

now we're going to talk about

reinforcement learning which is

completely different. Um it's it's

completely different than the other

three. It's a type of machine learning

where we uh basically learn from

interaction with the environment. And

you might ask what are we learning? We

are learning what actions to take in the

environment. Um and the way we do that

is by reinforcing

positive actions that lead to a a

reward. Um, so that's where the word

reinforcement comes from is we we

basically uh imagine like a child

that's, you know, learning from trial

and error. Like they're trying to crawl,

they're trying to walk and they keep

falling down. um eventually they learn

how to do it through trial and error and

they might get a reward

or they might um reinforce some of those

positive movements that lead them to

walk or crawl um or they might learn

from the penalties, right? They might

learn from uh some type of feedback. So

they might learn from falling down like,

"Oh, that hurts. I should uh support

myself a little bit better, right?" Or

be a little more coordinated. Um

and so they they learn from those

actions and their interaction with the

environment. Um

uh so this is a complex um algorithm

essentially uh it's it deals a lot with

um again taking actions. Usually when

you take an action something changes in

the environment um then you kind of

observe some type of feedback. So, think

about like a a board game where you're

trying to figure out what move you

should make. Or another good example is

like with a robot um trying to navigate

a maze. So, like what route should it

take? Should it move forward? Should it

move backward? Should it move left or

right? Those are different actions it

can take. Also, like a self-driving car,

should it should it turn? Should it

speed up? Should it slow down? Those are

all good examples of things that have

been trained from reinforcement

learning.

Uh yeah. So real world examples would be

like in a board game, uh a a reward

would be like if you win the game. Um or

if you like capture a piece like in

checkers or chess, that's a reward. A

penalty would be like if you lose the

game or lose one of your pieces, that

could be a a penalty.

um in a board game or sorry in like a a

robot navigation task, it could get

rewards for um moving in the right

direction

um towards the exit or like when it like

let's say you wanted to train a robot on

how to open the door and navigate a

room. Um you would penalize it for

bumping into the wall.

Um you would give it a reward for moving

usually oh like oh the algorithm

themselves usually it's like a a step

function um it's usually it's like a

discrete function that kind of is based

on the state so the reward it could be

like um like depending on the let's

let's go back to the board game example

like the reward could be like or even

the maze let's say like a navigating the

maze like getting to this let's say this

was the exit

and this was the entrance.

Then if they make it to here, they get a

numerical like if they make it to the

exit, they get a numerical reward of

like plus 100, let's say. So it's just a

number. And then if they uh like if they

bump if they go into here, like let's

say this is kind of like a death trap or

like a pit, this this would be like a

minus 100. So it could be like discrete

numerical values could be the reward. If

they're moving in the right direction

like let's say we want to encourage

going this way then we could give

smaller intermediate rewards like this

should be a plus like if you move

forward this is a plus five this is a

plus 10 this is a plus 15 if you're

moving in the wrong direction away from

the exit. Um that would be like a minus5

or a minus 10. Does that make sense? So

they're they're numerical in nature and

what you're trying to do is collect the

most reward. You're trying to get the

largest reward you can through trial and

error. So you you try this out many many

many times. You basically simulate

running through this maze many many

times. And what dictates it what

dictates like where I should go is based

on what I've observed in the past. It's

almost like you're a child remembering

like, okay, what move should I make from

this space? Like, if I'm here, if I'm

here, which way should I go? Should I go

down? Should I go right? Should I go

left? You kind of know that from

experience.

Does that make sense? Based on the

reward that I've seen in the past, like

when I've moved down, I've gotten a

higher reward than moving left or right.

Does that make sense? So, yeah, it's

it's a numerical value

as a reward.

Yeah, that's a great question. Um, how

does it differentiate rewards based on

gain and loss i.e. chess? So it's it's a

very comp complicated uh answer but

essentially every so in the chess board

you can think of the board as like every

every um

space is a state.

So I could be in this state I could be

in this state and then it's not not only

is every every uh space but where all

the other pieces are. So there's lots of

states that are possible.

Um, so

the way there's a way to quantify

essentially what's the value of taking a

certain action like moving my piece

left, moving it right, moving it up or

down um given the rest of the state. So

you're you're right, it may be

beneficial to sacrifice. Um, but we

would learn that through experience that

okay, the best move in this situation is

to sacrifice.

We would we would have to learn that

through trial and error many many many

times which is to say like okay if I'm

in this current state of the world right

all these pieces are distributed in this

way the best move for me right now in

the long run to get the most reward in

the long run is to actually sacrifice my

piece and move it right move it into

like a bad position theoretically but we

know from experience that's actually the

most long-term reward is from that

position

like moving it right may be the best

action for me. So what you learn is how

to take actions

and actions are usually like move right,

move left, move up, move down. You think

about like a self-driving car though,

that's going to be like slow down, speed

up, turn your wheel 10ยฐ,

um those kind of actions.

So the the short answer is it's there's

a calculation there that you learn what

the long-term value of every state is

every unique state

and then you're trying to basically say

what action should I take from that

state given that current state of the

world.

Okay.

And I really I really like reinforcement

learning. It's actually probably my

favorite field of machine learning.

Unfortunately, we won't be covering it

um in our main uh course. We have

offered uh electives around

reinforcement learning in the past. So,

um stay tuned. Maybe when we get to the

end of this program, uh we'll offer an

elective on it and if enough people sign

up for it, we'll we'll run it. But, um

we it's not part of our we don't really

cover reinforcement learning as part of

our main topics. It's it is an advanced

uh more advanced topic than than what

we'll cover. But, um I I really enjoy

it. Find it very fascinating.

Okay. So, all of this is kind of um

illustrating what I was saying, which is

um you think of like uh the thing that's

interacting in the environment like the

robot or the car or the human moving a

chest piece is known as the agent. It's

interacting with the environment by

taking actions which updates the state

um of of the environment. So that's

that's why you see this word state here.

This gets updated constantly every time

you take an action. Um ultimately what

reinforcement learning is trying to do

is learn the best action like what would

be the best action to take. Um

and the best action is is the one that

leads to the most long-term reward.

That's the best action. Um, so you have

to uh you have to learn what you know

what leads to a good reward by kind of

experiencing this over and over and over

through trial and error. So there's a

lot of um kind of simulation or letting

the robot try something a lot um in

order to kind of learn what's rewarding

and what's not. Think about it again

like I think a good example is like with

children, right? you kind of have to let

them try things until they learn on

their own what's what can they do and

what can they not do

what's the best actions right

so reinforcement learning has made its

way into other places so I I said like a

good example is self-driving cars or ro

robotics a lot of reinforce

reinforcement learning is used there one

place it's found its way into recently

is recommendation systems have kind of

merged with reinforcement learning

learning. Um, and this is because you

you can imagine there's kind of a

built-in reward for you clicking on a

video and kind of watching it.

Um, so that kind of reinforces that

recommendation and then uh that's where

um you can then kind of recommend a

similar thing and see if that's

rewarding and generates a click or

generates some view time or watch time

or whatever. Um so reinforcement

learning has found its way into a lot of

areas. Um recommendations being one of

them because it's just natural for the

idea of like what um should I recommend

next to generate the most reward. In

this case the reward is kind of

correlated to did they click on it or

not or did they how long did they watch

for longer it's more rewarding.

um those kind of things but uh place

places where reinforcement learning have

been used I said self-driving cars um

games so uh one of the most famous

examples if you want to look it up is

the um Alph Go this was in 2016 um the

Alph Go uh algorithm was a reinforcement

learning bot that beat um some of the

world's best Go players which go if

you're not familiar Go is a um board

game

that is a little bit more uh complex

than chess. It has more more uh it's a

larger board um more pieces to it. Um

but they there was a reinforcement

learning powered bot that actually um

learned how to play the game so

effective it could beat um world kind of

masters at the games was pretty amazing.

Um that's the alpha go and that was by

deep mind Google and deep mind in 2016.

That was pretty that was only in 10

years ago not that long.

Um so certain uh we said recommendation

uh even autocorrect um learning to

predict like what is the best correction

uh to generate a reward which would be

like you accept that correction or you

reject it would be a penalty. Um so

reinforced learning has been adapted to

these kind of problems very

successfully. Let's take a look at the

packages that we will use throughout. So

um of course we will rely on these three

which we've already relied on to do a

lot of things like numpy to do numerical

manipulations and calculations.

Uh mapplot lib to do any plotting and

not only mapp but maybe seabour as well.

both of those to do plotting. Um, pandas

is a big one because

that's where all of our data is going to

be manipulated and prepped before it

goes into modeling.

So, all of that stuff we learned from

pandis is definitely going to be applied

here in this course uh as we actually

build models. Um, so of course like

these old ones that we've been working

with quite a bit um still going to be

useful here in the modeling stage. Um,

mainly for different reasons though,

mostly to get our data prepared to do

some type of modeling or maybe to

visualize it before we do modeling to

get a sense of what it looks like, those

kind of things.

Um,

sci is sometimes useful for certain uh

um processing like in unsupervised

learning. We'll actually use scyp a

little bit to do dimensionality

reduction or help us do that. Um so

scypi will be used here and there and

we've seen it before with hypothesis

testing we use scypi like the test and z

test came from there. Um some of the

unsupervised learning stuff will come

out of there but the package we will use

by far the most in this course is going

to be scikitlearn

which is here. Um and we've already seen

a little bit about scikitlearn in terms

of its pre-processing capability. So we

use the uh minmax scaler and the

standard scaler from there from the

pre-processing module in scikitlearn.

But it has um many different models

built into it that we can use to help uh

do our training and predictions. Um so

it's a incredibly useful machine

learning library. It is the industry

standard machine learning library. Um if

you're going to do anything in machine

learning, it would be expected that you

know how to use scikitlearn.

Now what's really lucky about that is

that scikitlearn is a really easy

package to get used to. Nearly

everything we do in scikitlearn will

mostly follow the same pattern and so um

the code will be extremely simple. They

did a great job with that package of

making things really user friendly,

really simple. Um, it's a really

fantastic package and we're going to get

a lot of practice with it uh as we go

along. Every model we build will

essentially be from scikitlearn

and not only like the models but um

doing the training, doing the

predictions and then doing the

evaluation will all come from different

uh scikitlearn u modules. So that'll be

really nice and we'll get um good

exposure to that package throughout the

course. So if anything will come away

from this course as um scikitlearn uh uh

experts that'll be very nice. So this is

this will be the new one for us learn

but we'll get a lot of practice with it.

Okay.

All right. So just to recap that lesson

before we move on to lesson three. Um we

talked about machine learning as

learning from data um which is included

underneath the AI umbrella but deep

learning is also included under machine

learning because it's still learning

from data but it's learning using neural

networks.

Um we talked about the four different

types of machine learning. We had

supervised, unsupervised,

semi-supervised and reinforcement. So

those are the the different types of

machine learning that are out there. Um

and then we talked about some of the pi

Python packages uh that we will use. The

main one being scikitlearn and of course

we'll use our older like pandas to

manipulate our data and get uh pass it

into our model training etc. But

scikitlearn will be uh our goto for

anything machine learning.

All right. So, I have some questions for

you guys, some checks.

So, let me know in the chat. What do you

guys think? Uh, which of the following

best describes machine learning?

Which choice do you think makes the best

is the best for this?

Very good. Very good. I see I see a lot

of choices for A and A would be the

correct choice. So machine learning is

definitely um a a subset of AI. It's

underneath that AI umbrella, but of

course we're learning from experience

and of course that experience is

recorded in the data um without being

explicitly programmed. Uh so it's the

exact opposite of BNC. We're definitely

not learning from rules and it's

definitely not just used for image and

speech speech recognition. It can be

used for many other things beyond those.

So yeah, A is the best choice there.

What we say here?

Okay. What do you guys think about this?

Which example illustrates the use of

machine learning to enhance customer

experience in an ecommerce company?

In other words, what would be some what

would be some uh typical use cases of

machine learning?

Good. So I think uh C is going to be the

best answer here. Definitely C. So it's

using machine learning to do uh fraud

transactions. So so that would be a

prediction probably a supervised

learning, right? If if this is fraud or

not fraud. Um, and then maybe some

customer behavior uh that might be

unsupervised. So maybe grouping together

customers uh clustering them based on

their data like their shopping behavior

and characteristics. Um that that might

be unsupervised but either way it's

machine learning.

Okay.

Okay. Final one. What distinguishes deep

learning from machine learning in

artificial intelligence? So what's

unique about deep learning?

Oh, very good. Yep. So, deep learning

uses neural networks as so you guys are

right on top of that. Neural deep

learning uses neural nets. That's what

makes it unique. So, machine learning

would be part A. Machine learning is

focused on learning from data.

Underneath of that is learning from data

using neural networks which is what uh

deep learning is.

Very good.

All right. Let's go to lesson three.

And lesson 3 has two notebooks. We're

going to be starting with 3.1.

So, you'll want to open up that

notebook. I'm going to go over to it

now. Give you a moment to open that up.

So, we're going to open the 3.1

notebook. Um, there's two of them. We'll

see how far if we can get into the

second one today. probably will.

Um, but we're going to do the uh we're

going to start with 3.1 notebook. Do you

guys have this notebook? Should be in

your materials for for this course.

Let me give you a moment to open that

one.

Do you guys have it?

Okay.

All right. So, we're going to start by

talking about uh supervised learning.

um in our machine learning journey. So

remember we're going to talk about uh

supervised and unsupervised after we do

supervised

um and there's going to be a lot to

cover with supervised mainly because um

there are uh two different types of

problems we can tackle uh which will be

uh we'll talk about in a moment

predicting different kinds of values. Um

but let's talk about the kind of what

we're hoping to learn here which is um

talk about the different kinds of

problems that we'll study which are

these these categories of supervised

learning. Um those two categories are

going to be called classification and

regression. We'll talk about those and

their differences and then talk about

some applications and some uh example

algorithms

and that's just within this notebook. Um

3.2 two we'll get into uh regression in

particular um which will be uh very very

interesting. Okay. So that'll be our

first models that we'll build will be

over there in 3.2.

Okay. So if you guys remember um

supervised learning is where we learn

from labeled data. So we have input and

outputs in our in our data set. Um and

so you train a model on this data that

includes input features and

corresponding outputs that are that are

the labels, right? So um the goal is to

learn a relationship between the input

and the output. Of course, that's what

any model is trying to do. Um, and what

this allows us to do is then take that

model and use it to make predictions on

never-beforeseen

uh data. Right? So then we have a

predictive model out of that that we can

use um going forward on new examples. Um

so

remember we will have in our data a

bunch of features which are columns and

then generally one of those columns will

be the label that we're trying to

predict.

And our model is going to try to learn

some type of relationship between those

inputs and the output label. So the

output label could be like fraud not

fraud, cancer, not cancer, uh a price, a

temperature, those kind of things.

So let's talk about that. inside of um

supervised learning there are two

different types of learning that we can

do and they're really based on the label

or sometimes that label is known as the

target that we're trying to predict. Um

and depending on that type we get these

two different categories of learning or

two different types of learning. One is

known as regression. So that's generally

when we are predicting something that is

continuous or something that is a

numerical.

So numerical

numerical value. So think of price,

think of temperature, think of revenue.

We're trying to predict something like

that. Um versus something that is

categorical. So that the predicting

something categorical would be like

fraud, not fraud, spam, not spam. um

those are discrete categories and the

problem of predicting categories is is

known as classification because we're

trying to classify examples as belonging

to one category or another.

So we have these two main types of

supervised learning problems. we have

regression and we have classification

and they're going to be handled slightly

differently um for many reasons that

we're going to uncover. Um one of the

primary reasons is that of course we're

predicting something that's continuous

in the regression case versus something

discreet. So the models have to be

slightly different to account for that.

Um but then a step beyond that is the

evaluation has to be different too. Um I

kind of alluded to this last week, but

when you're predicting a regression,

it's very very difficult to to get the

exact numerical answer. So um generally

we don't care about that. Um generally

we don't care about getting it exactly

uh we don't care about getting it

exactly right.

um we just care about getting it um

we're just we care about getting it

nearby, getting it close enough. Um

whereas classification, we do care about

getting it exactly right because it's a

discrete category. So we're going to be

able to evaluate that a little bit

differently to say did we get the answer

right or wrong. Regression is going to

be did we get close? Um because it's we

assume it's going to be nearly

impossible to predict a continuous

number. Um, that's very hard to do.

Okay.

So, any questions on

uh that?

Any questions on those two differences?

Let me give you some examples. Maybe

it'll it'll help too.

So, again, the classification is going

to be predicting uh something that's

categorical. regression is going to be

predicting something that is continuous.

So think about trying to predict the

price of a house based on those other

features we talked about before like

square footage, bedrooms, bathrooms, all

of those things we predict the price.

That would be a regression problem

because the price is a continuous value.

Let's take a look at an example here.

Um, imagine we were trying to uh predict

the temperature tomorrow. That's going

to be a regression problem, a a

supervised learning kind of regression

problem because we're trying to predict

a numerical temperature.

Okay? And versus a category like a

discrete category would be this would be

a classification. So this is a

regression on the left. This is a

classification

on the right. Classification

um because we are um predicting one of

two categories. Is it just hot or cold?

Now, we're not saying exactly where that

threshold is on what's hot or cold. That

would be a decision on on what we want

to what our discrete categories actually

mean.

But, um we only have two choices, hot or

cold.

versus predicting the entire temperature

which would be um a numerical prediction

of some exact number. Right? So that'd

be a regression and then on the right

would be a classification.

Um now again why is this so different?

You can see the types of predictions

we're making are completely different.

One's a number, one's a category. But

again with the valuation it's like if

the if the true answer in our labels was

84

and we predicted 83 that's a pretty good

result. That's still pretty close.

That's pretty close to this. So from an

evaluation perspective that's pretty

good. Um whereas like if I predicted

cold and it's actually hot that's that's

a wrong answer. So they're evaluated

slightly different.

Um, and that's something we're going to

see as we talk about evaluation of our

models once we build them is depending

on if it's classification regression,

there's going to be different ways of

evaluating them.

You can kind of see why it's very

difficult to say, okay, we got exactly

84 when it could be any number. Our

model is going to be predicting a

number. That's really hard to pin down

an exact floatingoint number. So, the

best we can do is kind of say, how close

did I get? Like, this would be a worse

answer. If I got something all the way

down here, that's a really long distance

to here. That's bad. That's a bad

prediction. But if I get something

really close, that's better, right?

That's a decent prediction because it's

pretty close,

right?

Of course, being perfect would be

getting exactly right, but that would be

nearly impossible to do.

Okay.

All right. Any questions on this?

Does it make sense on regression versus

classification? We're going to use those

words quite a bit as we go along. So,

regression predicting that continuous

value. Classification predicting a

category.

And they're going to be um different

models that do that

different models being used for

regression versus different models being

used for classification.

All right, let's talk about supervised

learning uh applications here. So just

to name a few, we have HR operations. is

imaginary recruiter tasked with finding

the best candidates. Um so supervised

learning can help by um rejecting or

accepting candidates. Now this is

something that happens quite a bit even

today. Um and that it's kind of like uh

how recommendations happen like this

this resume should be um recommended

this should not um from a whole pool of

applications. Um so there's those kind

of use cases of of um predicting a

category that would be like a

classification. Should we should we

accept or reject the the candidate?

Um finance you see this all the time

with things like risk and loan

approvals.

Um you can uh predict the the the

category of like if the if the loan if

we should accept or reject the loan

application.

um you know that would be a

classification.

Um what's interesting about

classifications by the way so it says

here like we can predict the likelihood

of a of a loan being repaid

um is a lot of classifications um we we

say that they predict a category but

under the hood they can actually predict

a probability and we turn that

probability into a category. So, um, you

know, like we could say what's we could

say the likelihood of her loan being

repaid is very low. Let's say it's less

than 50% probability. Um, then we could

label this as reject,

right? We could label that as a

rejection. Um, if it's greater than 50%.

Then we could label this as accept. So

we can set a threshold there

and say okay truly we're predicting a

prob like our model spits out a

probability but we turn that into a

category by saying should we accept if

it's less than 50% we should reject if

it's greater than we should accept.

Okay, so that's something we will see

with some of our classification models

is that they actually produce a

probability and we turn that probability

into a category label

um by by doing something simple like

this putting a threshold on it um for

the for the category.

So finances is used all over the place.

Not only just loans like fraud, we

talked about fraud, not fraud. That

would be a classification.

Um predicting sales revenue, that would

be a regression, right? What is the

revenue going to be in the next two

quarters? That's going to be a

regression problem.

Uh emails like spam, not spam, that's

going to be a classification.

um that's going to operate on the that's

going to take the text input and predict

if this email is a spam or not spam.

That's going to be a uh supervised

learning problem, but it's going to be a

classification problem,

right? Uh manufacturing supervised

learning is used to inspect and uh

quality and classify products in

different grades. For example, a factory

might use a model to check for defects.

So this actually something that happens

is you look at images of products as

they go through the assembly line and

you can take a look at those images and

predict if it's a high quality, low

quality, medium quality. Um so they can

be this is a classification, right?

They're going into different categories

of quality. Um so it's much much like a

manual kind of intervention by some uh

QA or quality control uh specialist.

Okay. But that's a classification.

So in the maritime industry, supervised

learning can be used to predict current,

so current level

um and that can be used to forecast uh

supply and demand. Um so those would be

like regression models that are used to

predict um kind of like temperature, but

in this case like title levels.

We talked about fraud already, so that's

there. Um, that would be a

classification.

Okay,

any questions on these uh examples?

Of course, there's many more. Um

recommendation is kind of like a

supervised learning problem uh where you

are

taking examples of things that people

have viewed in the past or or reviewed

in the past and using that to predict

what they would want to watch in the

future. Um so recommendation is

supervised learning. Um and it's like a

classification, you know, trying to

predict um uh certain number of

categories of of uh shows or movies that

you would want to watch. Um

and that's something that we will study

in the future. Recommend we'll we'll

have a whole lesson dedicated to

recommendation as well.

All right.

So when it comes down to the uh actual

models themselves, so there's going to

be lots of different models that we are

going to cover. Um and they are um going

to be different in their purpose and

kind of their uh what kinds of problems

they're used for. Um and uh their their

how they actually train is going to be

different. Um, but at a high level,

they're all trying to do the same thing,

which is learn some sort of relationship

between the input data and the and the

label, right? That's really what they're

trying to do because they're all

supervised. They're they have those

labels, trying to build some

relationship there. Um, they just do it

differently.

And what we're going to study is the

pros and cons of a lot of these models,

like when would I use one of them, when

would I use another. Um, so we'll try to

talk about that as we go along. Um, but

they're all trying to learn some

relationship between the input features

and the output, right? So we have to

keep that in mind. They're trying to

model that relationship. They just do it

in different ways. Okay? So as we go

along and learn about new models, um, we

will learn the details. will learn the

ins and outs um and those pros and cons,

but they're no matter what, they're all

trying to

uh learn that relationship, right? And

be able to make predictions on new data.

Okay,

so here's a list of models that we will

cover and work on throughout the uh the

sessions that we have.

um we're not going to do them all in one

one sitting, but um the first one that

we're going to start with and that we'll

cover today is going to be linear

regression.

So we will cover linear regression and

then we'll cover the rest of these guys

mostly in the context of uh

classification.

So, um, what's interesting is some of

these guys can actually be used for both

regression and classification as long as

you make, um, certain adjustments to

them. They have variations that can be

used to do classification and regression

is very interesting. Um, but we're going

to start with linear regression today

and then work our way through the rest

of these models when we do um, we're

going to do a separate lesson four on

classification. And so these all these

guys will come from lesson four.

Um and then uh we will do this guy in

lesson three in the 3.2 notebook. We'll

do all about linear regression.

Yeah. I so logistic regression is a

classification. Um which is kind of

strange that its name is regression but

it's doing a classification. But the the

reason is that the logistic regression

um computes a probability. So it does a

regression to predict a number but that

number is actually a probability. So it

it produces a result that's between it

produces a probability that's between um

obviously uh zero and one.

So it uh and then we take that

probability and we turn it into a

category

like a spam not spam fraud not fraud.

Um but so so logistic regression is kind

of special. It's sort of like a

regression but it's predicting a very

specific type of value which is a

probability. So for for that reason it's

a classification uh algorithm primarily.

So we'll study that one in lesson four.

Uh but yeah, that's that's why it's

under that kind of umbrella of

classification is because it's it's

producing a probability as its main

output which we can then turn into a

category as long as we interpret that

probability as um in the right way. Uh

like the probability of spam,

probability of not spam.

Okay.

Okay. So, let me focus on um

let me focus on linear regression. I'm

not going to go through all of these

other use cases because we haven't

learned these models yet. Um so, I don't

think they're good. Uh I don't think

it's good to read about them yet until

we've covered them. So, once we cover

them in lesson four, I'll come back and

describe these examples to you guys and

we'll see why it makes sense. But I

think for linear regression um which is

what we'll cover next, let me talk about

that example. So a prototypical example

would be like predicting the house

prices that we've seen in that house

price data set.

So um if we wanted to uh if we wanted to

predict um if we wanted to estimate the

market value of a house so the price

um we could do that by using the

features such as number of bedrooms,

square footage, location, age of the

property. Um and you know then when a

new when a new house comes on the market

we could estimate what the price should

be based on those features. So linear

regression is a good one to predict the

price like a housing price. Um and we'll

actually practice that in the next uh

notebook.

So we'll we'll uh and then all these

other now there's descriptions of these

other models but again we haven't

covered these guys yet. So I don't want

to really go through those until we get

to those models. So we get to those I'll

come back and mention the example.

Uh can K andN be used for clustering?

No. So um the clustering model is going

to be different. It's going to be uh K

means

K means that's the primary clustering

model. Not K nearest neighbors. K

nearest neighbors is used for uh it can

be used for regression. It can be used

for classification.

So we'll we'll talk about K andN which

is the K nearest neighbors in lesson

four.

It sounds really similar. Yeah, it

sounds really similar but K means is a

clustering algorithm that's that's

slightly different

different uh there's no labels used at

all. This K nearest neighbors is a is a

supervised learning algorithm. It uses

uh labels.

Good. Any any other questions so far?

Okay.

So that being said, let's move on to the

3.2 notebook.

Let's move on to that which will be our

um first discussion around uh

regression. So going into supervised

learning and regression. Give you guys a

moment to pull up this notebook.

But yeah, you want to pull up the 3.2.

We'll do this one next. So we'll focus

in. So our plan is to do regression

first and then we'll talk about

classification in lesson four

which we will cover all those other

models which you you could use for

classification uh on that list. But then

we're going to talk about linear

regression uh first.

All right. So we have a a big agenda.

This is a big notebook um to go through

a lot of material here surrounding

regression. So we're we're going to

start with linear regression and see um

how we actually perform it, what that

model is doing. Um which we've kind of

seen the idea of it a little bit

already, so it should be somewhat

familiar. Um and then we'll talk about

how to adapt that linear regression idea

to um nonlinear what's called nonlinear

regression which is going to be using

like polomial uh features. We'll talk

about how to do that. Um and then a big

big big topic for us is going to be

evaluating the model. So it'll be it'll

be quite easy to actually build it.

building the model will be really easy

but evaluating and interpreting that

will be uh a lot of interesting work

there um because we want to know what

the performance of that model is once we

have it built right we want to know how

good of a model is it is it worth using

or do we need to retrain it or get new

data or change the model up to talk

about that um how do you determine what

to do based on that performance

um and then we'll We'll talk about here

um a couple things. We may not get to

this today, but regularization

which is used to boost the performance

uh in certain situations um whenever the

model is kind of performing um poorly

against test data even though it

performs pretty well on training data.

In that scenario, you can use offshoots

of linear regression that do some uh

what's called regularization. We'll talk

about that.

Um, and then we'll talk about

hyperparameter tuning, uh, generally as

a strategy, which is something you

generally do want to do when you're

training machine learning models. Um, so

again, these two we may not get to

today, but um quite a quite a lot to get

to be prior to that mainly centered

around evaluation and building linear

regression.

Okay, so pretty cool. we'll get to our

first kind of model here. This linear

regression

to start with.

Okay,

so let's start with uh linear regression

here. Um, and really what linear

regression is attempting to do and I

want to show you this in this picture is

draw this line sometimes what is known

as the line of best fit. So this is our

model that kind of goes through the data

and it's generally a good predictor

um because if you give me um features uh

if you give me new features and let's

say they are let's say you give me a

feature that's right here.

So you say, okay, I have a feature

that's this value on the x- axis. Then I

know all I have to do is plug that into

my line equation, and I will generate a

a value that's like right here.

Okay, that's pretty that's on that line

at that input. And that's going to be my

prediction for what the output variable

should be. It's just going to be

something on that line. And what you can

see is this line is a decent estimate

for this data because it slices through

this pretty evenly. So it's a good guess

as to what the output should be given

any one of these inputs. It's a it's a

good estimator this line. And so our

goal building a linear regression is to

kind of build the equation of this line.

So we want this equation.

Equation of this line

is going to be our model.

Yes, it's going to look just like that.

MX plus B or yeah, MX plus C. It's going

to look exactly like that. uh except

that it's going to be more than just MX

because we have um generally more than

one feature. So you think of X as a

feature um it will be more than just MX.

It will generally be like uh it'll

generally look like this

and then plus maybe some bias here plus

an intercept. Yeah, it'll generally look

like that. So, yeah, you're exactly

right. MX plus B is the right idea.

Exactly right.

It'll generally look like that.

Nonlinear, it can be adapted to

nonlinear. Yeah. If we transform, we're

going to talk about that. If we

transform all of our features in a

nonlinear way, um we can apply linear

regression to it. Yes. And and that

would be a nonlinear regression. So yes,

we can do nonlinear things too.

We'll talk about that.

Okay. So linear regression again is the

art or science I should say not really

art but it is an exact science of

finding the equation of this line that

fits through this data. Um now why one

thing you should be thinking about is

why is this line a good predictor and

the argument is that if you take a look

at this distance from these blue points

so let's say these blue points are

actual data points this line is going to

be found such that it minimizes this

distance

from the points to actually I should

draw it this way from the points to the

line. So, we want this distance to be um

actually I should draw it that way. This

way. We want this distance to be kind of

at a minimum. So, it would be bad to

draw a line all the way out here because

then that's a lot of distance, right?

So, and that would be a lot of error um

contributed from not being able to

predict those points in our data set

very well. Um which is our training

data. That's why we have labels, right?

that that guide us in building this

line. Um so our goal is to build that

line especially so that this error or

this distance can be as minimum as

possible. Right? Which are all these

distances from these points to the line.

We want those to be as minimum as

possible. So our goal is to find this

equation.

So we're going to build a model that's

going to find this equation.

of the line

um such that our error

is minimal.

And what is the error? The error is the

distance

of our data points

to to

the line that we build. So essentially

what we'll do in order to train this

will be to adjust the parameters or the

or in that like I think is really good

you brought up the MX plus C. Basically

the M and the C will adjust. So we

adjust those accordingly to make this

distance as small as possible.

Okay? To minimize that distance as much

as possible.

Okay.

So um where is regression used? We've

already seen some examples. Here's some

more uh advertising like predicting

sales, predicting um oil and uh oil

production and demand. Those are like

forecast those are regression problems.

Um retail like demand forecasting for

inventory. Um healthcare predicting um

uh the levels of certain um uh blood

markers or you know something like that.

um real estate predicting prices based

on those uh talked about like square

footage, bedrooms, bathrooms, those

things. So regression is used again

whenever we want to predict a number a

numerical output um that's a regression

problem.

So this kind of regression we're talking

about here is generally

um known as uh a when that equation is

linear that is known as a linear

regression. So go back to that picture

when we have a when that equation of the

line that we find is a linear equation

meaning that it is exactly the form I've

been telling you. So it's it's something

like um weight time feature

plus weight time feature

plus weight time feature

and then maybe some intercept um term

like some some bias term there.

Um this is a linear equation because all

of the features are to the single power.

So it's a linear power and this is a

linear combination of features with with

those different weights. So this is a

linear model

because it is uh it's what in math we

would call this a linear equation right

everything is to the first power. It

resembles mx plus b. It is a linear

equation or linear model. Um so when we

talk about linear regression that is a

regression model so we're predicting

some continuous target that assumes we

are model our model is formed from this

kind of equation a linear equation.

So this is going to be our our model for

a linear

uh regression.

Okay.

And so when you when you train a linear

regression, your goal is to learn these

weights so that you can plug in um you

can plug in any one of your uh input

features and you um can generate a

prediction. You can which is going to be

something on that line, right? It's

going to be a value that's sitting here

on this line.

We put in all of our features and we end

up there somewhere on that line.

This output.

Okay.

Okay. Let me pause there. Any questions

on the linear model here or why it's

called linear regression?

Okay. And by the way in these notes um

this bullet point here where it says it

uses the least squares criterion to

estimate the coefficients that is

exactly what I said earlier with the

distance. So the distance is based on

the square

of this this quantity like how far away

you are from the line is based on this

square distance here and here and here

and here. So what we're trying to do is

find the least distance or least squares

which is that minimum distance. So

that's how we find all of these weights

is from minimize. We basically tune them

enough using our labels. So here's our

label which is the y. We basically plug

in our data and tune those enough to

minimize the error. It's it's a it's an

optimization problem, right? We we're

trying to find the minimum of this

quantity which is that best fit line.

Okay.

So we have linear regression

um and we can do a simple linear

regression that only has one feature. So

if it only has one feature that's

exactly the so if there's only one input

feature sometimes that is known as um

simple regression or simple linear

regression and there's only basically

there's only one feature. So one

independent variable is the feature.

There's only one feature. And so this

equation resembles the

exact equation that you guys just put in

there, which is um mx plus b,

right? It resembles exactly that. um

we're just using different symbols for

those like beta beta 0 and beta 1 but um

basically exactly that simple line

there's only one feature. So and that's

because that line is going to um that

line is going to be generated uh

according to that equation. So here's

kind of what it looks like.

This is the best fit line through all of

these blue dots. This is something we're

going to be able to build. we're going

to be able to build that equation um

pretty easily in scikitlearn.

So we'll be able to find that um and it

won't be too hard. So this line will be

um y = beta 0 plus beta 1. So some

weight beta 1 times the only feature we

have x1.

Okay. So in this case um we would be

predicting sales. So sales would be the

value basically the label that we're

trying to predict and the feature that

we're putting in is uh I think it's the

number of TV expenses. Yep. TV expenses

which is on the x- axis. So there's one

feature which is um TV expense.

So um on this graph this would be this

would be our model.

Okay that would be our model. We only

have one feature and we have um these

two weights. We have an intercept B 0

and or beta 0 and then a one weight

which gets applied to that one feature

beta 1. And so our model would have

certain value for beta 0 and a certain

value for beta 1. That's what get that's

these guys get learned

learned during

model

training.

Okay. So those are what get learned

during our model training and they get

learned by a a a least what's called a

lease squares algorithm that is trying

to minimize that distance. It tries to

tweak beta 0 beta 1 to minimize this

distance of this line

um this line

to all of these points

trying to minimize this.

So imagine taking a line and kind of

moving it around and turning its its

slope, its angle um to try to find that

best fit,

which reduces that error the most.

Right? That's kind of what we're doing.

Uh can I explain? Yeah. So uh sales is

in dollars and and TV expense

um

uh

TV actually I think it's the other way

around. I think the sales is actually a

quantity. So I this is number of sales

that we have and TV expense is um I

think I think it's in dollars. So how

much money how much expense um did we

put into the into the product and then

this is how many sales did we have of

that product.

So I think it's the other way around.

But what this what this graph is showing

is the blue points are our actual data

points. Okay. So so we have a collection

like we have a data frame that has so

imagine we had a data frame that has the

uh true values.

So it has the um TV expenses. Um, it has

points that are like one. So, it has

points that are like 120 and then the

sale sales could be like 700

700 units, let's say. And then it has um

so this is just our data set, right?

This would be like in a data frame that

we have. And then we had ones that were

um 50 and then this could be um this

could be 400, let's say. And on and on

and on, right? So this is our data and

this data is plotted in the blue. So

these are these blue points here,

right? So these are the blue points here

and the red points are is our model. So

we built a linear regression model

um where we are putting in some values.

We're putting in some e fake x values

here and generating some predictions

which is this line

this linear uh regression line.

Right? And that line is derived from

this data. Right? It gets learned from

this supervised uh examples.

Does that make sense?

That line is derived from the data. It's

actually um learned from like the line

of best fit is learned from that data

and the actual data is in the blue.

So you can see we're trying to build

this such that this distance is kind of

a minimum.

So it's an optimal fit

to balance out these distances.

So it's just plotting. So it's just

building that relationship between the

input and output. like when the when the

expenses are higher, um we seem to have

more sales.

Uh what's perpendicular like the

distance? This should be this should be

perpendicular because it's a distance

here.

Is that what you mean? Like the distance

from the real points to the line. Yeah,

that should be perpendicular

because it's it's a it's a distance

formula.

Okay.

All right. So, more generally now do do

we usually have one feature? No. So

generally we expand this to the more

general case where we have more than one

feature like what we see in the housing

data right where we could predict a

price but we have many different inputs

like bedrooms, bathrooms, square footage

etc.

So more broadly

instead of simple linear regression we

have what's known as multiple linear

linear regression which means we have

multiple variables or multiple features.

Um so this is exactly the equation I've

been talking about. Um so we just extend

that that one into many features. So

which is this case and then a intercept

term which is uh um there as sometimes

known as the bias. Um

but this is the intercept term to kind

of orient the line to start out in the

right place. Um and uh but this is the

um this is the equation that we would be

building the model. This is our model

essentially, right? This is the equation

we would be learning.

Intercept is like a constant. Yeah. So

if if all of the features were zero, um

this is what our our data would be. This

is what our result would be. If

basically if this was zero, this was

zero, this was zero, it would reduce to

this as the prediction. Yeah. It's like

a constant. Yes.

So in in geometry, the intercept is

actually really important because it it

orients where your line should start. So

it orients like so so these values are

kind of like the slope. They orient the

tilt of it. Like should it be tilted

like this or should it be more sloped?

But the intercept orients where it

should start like vertically like should

it start all the way up here? Should it

start more down here?

Um, that's what the intercept kind of

tells us.

Okay, so this is the situation. This is

going to be our linear regression model

that we will be building most of the

time because we will have again these

are all going to be features.

So this is some feature the X this is

some feature this is some feature

X1 etc. These are all features and what

gets learned during the training are

these coefficients. So all of these

coefficients including the beta 0ero um

will get learned. So these will get

learned

um from our data right they get learned

they will be trained from our data um in

order and and how do they get trained

it's from reducing that distance we try

to get that line of best fit by tweaking

those betas enough to uh until we reach

a minimum distance but there's there's

an algorithm behind that um that that

scikitlearn will run for us to find that

best fit Um, so we don't need to do that

manually, but that's that's the process

is basically tweaking those weights to

end up with that line of best fit. So in

higher dimensions, instead of a line,

you get more of what's called a plane

here. Um, which kind of looks like this.

So the best fit is actually this plane

where all um, it kind of dissects all

these points just like that um, in

higher dimensions. So this is uh instead

of a line you get this in in three

dimensions you get this plane like this

but it's still it's like a line of best

it's just a more general line of best

fit. It's still the same idea. Um we're

still trying to um come up with the best

coefficients to minimize that distance

from our from our points to the line.

Although in higher dimensions it's no

longer a line. It's more like a plane

like this. So you're trying to minimize

this distance from here down to the

plane

here up to the plane

in higher dimensions. So I want you to

keep in mind what we're trying to do

before we go into the code because the

code's going to make it seem really

really simple and that's because

scikitlearn is great and that's what it

does.

But we should realize that there's

something really complex going on which

is again finding the best value of these

weights

that minimizes the distance of this line

to the data points that we have. So

there's an algorithm there that will

keep trying to make adjustments to this

based on those distances. So it's going

to use those distances as a guide to

kind of tweak them to find the one that

results in the lowest amount of

distance. So we keep making tweaks, keep

making tweaks, keep making tweaks and

eventually we try to find we converge to

the set of weights that gives us that

best fitting line. Um and and there's an

algorithm there that occurs. Now luckily

that gets abstracted for us a bit behind

um scikitlearn

um finding that best fit. So there'll be

a function that we use in scikitlearn

when we build the model that will go

ahead and find the best weights for us

and that's then we now have our optimal

model right that then we can just plug

in different values of these features

and generate a prediction which is going

to be this uh result right so so that's

what we're ultimately trying to do is uh

train the model which will uh find all

those optimal weights and then uh we can

predict with it which would be plugging

in different feature values to to

generate a prediction.

Okay,

so let's see how that happens. It's

actually going to be super easy um with

scikitlearn.

So uh in this scenario we have um we're

going to import our pandas because we're

going to load our data from that. Um, so

of course we need some data to work

with. So we're going to load this uh

CSV.

Um, I

uh so I was not actually able to find

this CSV for this example, but I mean

that's okay because we'll do some we'll

do other examples where we'll work with

the data. If you happen to have it, um,

great. I didn't see it in in my files.

So just have to take the word for it

that these are the this is that TV and

sales columns here um from this data

set.

Okay. Um as an example. So um just to

see how it's fit um what we're going to

do and this is going to be a very

standard process for us for building a

model. These steps are going to be very

very standard for us which is going to

be first of all splitting the features

away from the label. That's the first

step that we always will take. So if you

take a look at this code, it's taking

all rows but only the first column.

Okay, so it's extracting all the

features from the dataf frame um which

happen to be which is just the first the

first column uh which is the TV uh

column right just that column there and

our target variable which is our label.

So our target variable aka the label um

is the second column, right? It's that

that sales column.

Um and so our first step here, let me

call that out here. First step is to

always split apart

features from labels.

Okay, so we put all those features into

a data frame called X and we have all of

our labels into technically a series but

uh sort of like a data frame, right? Um

called Y, which is just the um which is

just the uh uh labels. So that's just

the TV values. Um now you're going to

see why we do that. It's because we need

um our our features and labels split

apart to put them into the model

building function. It expects our

independent variables or our features to

be separated from our answers or our

labels that guide the model building.

That's the first thing you got to do is

separate those.

Okay, so this code will separate those

out into a capital X and a lowercase Y.

And that's actually pretty industry

standard notation. Whenever you split

apart all your features, usually you put

them into a data frame called capital X

and then you have a lowercase Y to

represent your labels. That's actually

pretty standard.

So it's pretty standard that um X

represents

features

and

Y represents labels

label column

whatever our label column is in this

case it is the sales because we're going

to be predicting sales

using the TV column the TV quant uh

expense quantity.

Yeah. So what it so the assignment is

that we are um the assignment is that we

are

uh we are um splitting apart our data.

So that when we first read in the data

um it is a data frame right that has two

columns TV and sales.

Oh perfect thank you Tim. I will I will

go ahead and so if we look at this data

it only has those two columns right it

only has those two columns. Okay. So

what we're doing with this is we are

splitting apart

our our independent variable our

features. So this this X will contain

our features

and Y will contain

our label.

Does that make sense? We're splitting

this data apart. So, we're only grabbing

that first column here to be our

features. And then we're we're grabbing

the second column, which is the sales,

because we're going to predict the

sales. This is our label. We're going to

we're going to build a model to predict

the sales given the TV input, TV expense

input. So, the first thing we have to do

is split apart the features and the

label.

Okay, that's the first step we usually

will take. And the reason we have to do

that um just to reiterate the reason we

have to do that is because our model

will expect our our data features to be

separate from the label. We will pass

those in separately.

X is TV. It's the first column

because we're using right. It's the it's

all rows but the first column

which is TV.

Why is sales? This is what we're

predicting.

We are predicting the sales given the TV

expense value.

Yeah. Which is why we split it into So

this is the second column, right? The

index one column.

Uh you just put in read CSV and pass in

the URL.

So you could So exactly the code that

was up earlier from temp

um you just do this

and then data equals ddread CSV URL.

So we split our data into X and Y here.

All right. Now, one other step that

we're going to take that's a very very

critical step and you're gonna we're

going to see this step over and over and

over and over again. So, splitting apart

into X and Y will become we'll do that

over and over and over and over again.

Not only that, but doing this next step

which is what's called a train test

split. Now, let me show you what the

train test split does. It takes our data

And it's going to split apart our data

that we have, our X and our Y data. It's

going to split it apart into a

percentage that will be used to train

the data

and then a percentage that will be used

to test. Now, why would we want to do

that? It's mainly so we can do

evaluation. So we build the model over

here and then we test it on data that

has not seen before. So we reserve a

percentage of the data to be used for

test. Usually this this data is um

somewhere between uh 20 to 30%.

So somewhere between 20 to 30% of the

original data. So that means the

majority of it is used for training. So

the majority of the of that X and Y over

here is going to be between 70 to 80%.

Will generally be used for for uh for

training. Okay. So somewhere between 20

to 30 the industry standard is some

anywhere in between there. Um a lot of

people like to use 30%, some people like

to use 20%. Um anything in that range is

acceptable. um we will I think we

generally will favor like 30%.

Um to be used for testing but um the the

point is we don't we don't want to mix

those together. We want those to be

separated out so that we can have a fair

evaluation, right? We want to train our

data on this train our model on this

data and then see how well it performs

on this data that it has never seen

before.

Right? So in order to have data it's

never seen before, we're going to take

our X and our Y and we're going to split

it using this function called train test

split that will do this kind of

splitting for us. Okay. So scikitlearn

has a function called train test split

that will go ahead and we're going to

pass our x and our y and we'll pass in a

percentage like 30% that we want to

split out into a test set and then the

remainder of that the 70% will be used

for training the model.

Okay.

So what we're going to get let me redraw

that. So, what we're going to get out of

this for the train test split is we're

going to we're going to have an X and a

Y per

training and test. So, we're going to

get now we're going to get an X train

and a Y train.

So, we're going to get training features

and training labels. And then we're

going to get test features

to plug into our model and and test

answers or test labels

to do evaluation because what we should

be able to do is build the model over

here and then apply the model on this

data. Meaning we can take these features

and plug it into our model and then see

what answers we get and compare those

answers to this testing data. Right? We

should be able to do that to generate an

evaluation.

Okay? Now you may be wondering why do we

do any of that? What's the purpose of

that?

Evaluating it on this test data gives us

a good sense of will our model

generalize to new examples. Right? If it

performs pretty well on this data,

that's a good signal like when it's

performing pretty well on data it's

never seen before, that's a good

indicator that it's going to perform

pretty well when we use it on brand new

examples

um in the future.

Right. So that's a that's why we do this

evaluation on this data that it has not

seen before. It's going to see this

training data, right? We're going to

train the model on that data. But that

model will never be exposed to this test

data until we do the evaluation

and and generate some metrics to see how

good is this performing

and does it have a good chance of

generalizing to never before seen

examples which is what we want right

because we're going to use this model in

the real world. It's going to be being

used on new examples that it hasn't seen

before. We want it to perform well. So,

this is kind of our test, our

evaluation.

Okay. Any questions on the We're going

to do this in a moment. I'll show you

what it looks like in the code, but any

conceptually, any questions on the train

test split idea? It's a very very

important idea that we um basically use

part of the data to train it and then

another part of it to evaluate. It's

very important we do that. By the way,

this has a term um this in machine

learning this is called cross

validation

because we are using one data set to

train the model and then we're cross

over we're crossing that over into

another data set to validate it which is

the uh the the testing set.

So this is called cross validation. Um

there's actually many ways to do cross

validation. That's something we'll

study. This is a very simple way of

doing cross validation. There's more

complex ways. You can take your data and

you can actually divide it into many

sections

and basically train it against most of

these and evaluate it against one at a

time and then rotate. So that's another

way to do cross validation. We're going

to study that. Um but this is the this

is the simplest way to do it here.

Okay.

So let me show you what you get when you

use train test split. So uh we're going

to import from sklearn.

We're uh from the model selection

module. Now we haven't used this before.

This is our first time using it. But

here's our model selection. We're going

to import this train test split function

and we're going to use it on our X and Y

and we're going to set a test size of

30% which is which is.3. So our test

size

is 30%.

Converted to decimal

right converted to.3. So that means

we're reserving 30% for that test set.

Um you can set a random state. Now

that's completely optional. Um the

random state

is for reproducibility

because what the train test split is

going to do is it's actually going to

shuffle the data and then split it apart

into the 7030.

So um yes, the seed. Exactly. It's like

a seed. So it's it's saying like when

you do that shuffling every time I run

this notebook I'm going to get the same

result but it's going to be random the

first it's going to be random but I'm

gonna be able to reproduce that

randomness with that random state. Yes,

it is like a seed.

Uh it's you can choose any number to be

your your um your random state. It 42

isn't important. You could choose zero.

You could choose one. Um, you could

choose any positive integer. Um, 42 is

kind of like the uh industry standard.

It's it's you'd have to look it up why

it is. Um, apparently 42 is a special

number. Um,

in in kind of the history of development

of this stuff, there's nothing really

special about 42. You could choose a

random you could choose a random seed to

be uh zero. That's fine. It it doesn't

really it doesn't really matter.

Um you just want you could choose it to

be uh one, two, three. Um you could

choose it to be 15. You can choose it to

be anything you want it to be. It's

really so that your your shuffling is

consistent. Every time you run this

notebook, you get the same shuffle

result. So I'm always going to get the

same rows in these splits.

Hitch. There it is. I knew it was from

something.

Yeah. So 42 is kind of like a

it's it's just used ubiquitously

uh you know as kind of a um paying

tribute to the Hitchhiker's Guide to the

Galaxy, but it's no it's there's nothing

that special about 42. It doesn't it's

not going to change our result or

anything.

It's just so that this train set split

is going to shuffle our data and split

it apart into 7030.

You just want to set this to something

so that you get a cons every time we run

this notebook, we get a consistent

shuffle.

And so the data in these sets

are uh consistent. That's all.

Okay. But do you guys see how we pass in

our X and our Y and we generate four we

generate four different data uh

quantities here which is we generate

training features, test features,

training labels and test labels because

again we are generating these four

different we're generating data on these

two different sets. A training set and a

test set. So we have training features,

training label,

and then test features, test label.

Okay, that's why it's so important to

split apart our data into the X and the

Y. We need those split apart in order

for this part to work.

So by the way, these two steps we will

always do for any model we build. We'll

generally do X and Y and then train test

split in order to generate the data that

we will use for building our model.

Okay. So this this data here is going to

be what we actually use to guide the

training of our model. So it's

definitely supervised, right? Linear

regression

um we we will use that

Okay, so we haven't built the model yet.

We're just getting our data split apart

and ready for the training. We haven't

actually built our model yet, right?

That'll be coming up uh in a moment.

But this is getting our data ready. We

started with our data frame. We split it

apart into uh an x and a y. And we split

that into a train test split. And um you

know then we can uh then we can go ahead

and um pass in to our model training

which we'll do in a moment.

Um you that's a good question. You could

run so what you could do is you could

run

um should we import numpy? Let's see.

We did. Okay. You could run the average

on the um you could check the MP mean on

the X train and see how it compares to

um

see how it compares to X.

So you could you could do that and see

what the average of this feature is um

compared to the average of the original.

They may not be perfect because we are

taking a reduced data set size. So I

don't think there's really any good

there's not like a one-sizefits-all

validation we can do because we're

taking a random shuffle and taking a

percent. We're taking 70% of the data

out. So we're not guaranteed to maintain

the same statistics. We can see if

they're close.

Um but does that make sense? Like we're

not guaranteed to get the same stats

because we're taking a slice of it.

We're taking 70%.

So it's not guaranteed to to to

be the same distribution really.

Delete that.

Uh is it a good practice? Yes, it is.

It is. Uh 30% is the industry standard.

Anything between 20 to 30. So 0.2.25.3

any of those are acceptable. It's really

up to you. Um I mostly see 30%.

Mo I think.3 is is a good good practice

to use for sure.

Um I did explain random state. Uh random

state is so that you get consistent

shuffling. Um you can set this to any

integer that you want it to be. It it

doesn't really matter. Um you can set it

to uh 100, you can set it to 10, you can

set it to 15. Um it just ensures because

what this split will do is it will

shuffle the data first. It'll shuffle

the rows and then um split it apart into

the into the train and test sets. So you

set the random state so that the next

time you run this you get the same

consistent shuffling. That's the only

that's the only thing it it helps you

with because it is randomized but when

you set a random state um it's so that

like if you run it again you'll get the

same shuffling.

You'll get the same the shuffling

matters because it it it uh dictates

what ends up in in these sets.

Okay.

All right. So let's see let's do let's

build the model

um and let me show you how easy this is

going to be to build the model and this

is really how it's going to be for every

single scikitlearn model will basically

look the exact same for training it

which is what's going to make it really

really nice. So the first thing we have

to do is import our model. So from

scikitlearn we're going to be using a

linear from the linear model package or

the linear model module I should say

within sklearn we're going to be

importing the linear regression

and we're going to create an instance of

the linear regression here.

Okay, so linear regression and look how

easy this is going to be. Nearly all

nearly all sklearn models use

ffit function to train.

So every one of them, no matter which

one we use, like the decision tree, like

the um logistic regression, any of those

like we use for classification that are

going to be coming up in lesson four,

they're all going to look the same in

terms of it's going to run.

Which is um scikitlearn's

uh generic function for training your

model. So this will execute the training

once we run this code. And what that

again the linear regression training is

going to do that least squares distance

procedure or algorithm to try to find

the right weights. It's trying to find

those weights that minimize that squared

distance uh from our line that it's

trying to build to the data.

And what I want you to notice is what we

put into the ffit. See how we put in the

training data where we put in the

training features and we put in the

training labels. Now this is supervised.

So of course we put in the labels,

right? Of course we put in these labels

here and of course we put in our

features here. So we're putting in all

of our examples from our training split

into this ffit which is going to train

the model uh so that we can we can use

it for prediction.

Okay, it's really fast. If I run this,

it's going to be pretty much instant.

Pretty much instantly it gets trained.

And you can see here we now have a

linear regression. you can see in this

little box. Um, and it and this

information says that it has been

fitted. So, it's now ready to be used,

right? So, we now that's it. We've

trained our model. We tr That's how easy

that was. We did fit. Now, what we

should realize is there's a lot of work

going on behind the scenes of this ffit.

Okay, there's a lot of work being done

there to do the least squares algorithm

and find those weights and and create

that line of best fit. Right? So there

there's a lot of work being going on

there that's going on there behind the

scenes, but scikitlearn is abstracting

it away for us. Right? And all we have

to do is fit when we're using this code.

Really easy. Really easy. Fit. And there

we go. We've trained our linear

regression model.

And by the way, if you want to see what

the coefficients are, you can actually

extract them if you do so if you take

your lin regression and you do um

coefficients like this

coeff with a with an underscore. So this

gives us the trained

weights coefficients

also known as the coefficients right.

Um so if you run this you can see uh

right now we have this coefficient here

um which is the only coefficient we had

on our feature. So we only had one

feature coefficient there.

And we can take a look at our intercept

which is this.

So this gives us the train weights

and so we can look at the intercept we

can look at the the the coefficient. Um

so obviously if we have multiple

features our model has many features

it's going to have more values in that

coefficient but the intercept is just

the single value 7.23

and then the coefficient

is 0.046. So that's the weight that gets

learned.

Is there a size limit? No, not really.

There's no size limit. Um,

no, you can use as much data as you

want.

There's really no size limit other than

what like what you can fit in memory.

I'd say that's the only limit is

basically what the amount of data that

can fit in memory.

Okay.

All right. Were you guys able to run

this? Were you guys able to run the

linear regression ffit?

Okay, perfect.

Perfect. You are Okay, great. Great.

So, we have a model and we can use it to

predict. Um, and so that's actually what

we're going to do next. If we go down

here, um we're going to have a function

that's going to um build a scatter plot

of our original test data.

Um so we're going to have our test data

here.

Um,

and we're going to then take our uh

we're going to take our training data

and plot we're going to use the uh this

data versus our sales predictions. So

you can see we're going to you this is

how by the way this is how you use the

scikitlearn model to predict. You have a

fit to train it and look at the function

you use to predict. It's literally just

called predict. That's how easy it is.

and you pass in your data, all your

features into this predict and it

generates a prediction for every row. So

every row in these features in this data

frame um will end up with a prediction

using our model. So what we're going to

do is plot our training date uh features

against the predicted sales to see how

good of a fit that really was.

Okay. to see to see the regression fit.

Okay. And so there's the regression fit.

We have all of our test data here

plotted in the green. We have our blue,

which is our um we have our our blue,

which is our uh um training data line

that we built our model on. So that's a

pretty decent fit. Um, and then our test

data is here. We just plotted in the

green scatter. But the thing I want you

to see is this prediction, right? We we

were able to generate some predictions

on that training um by running our

predict function with our model. Now,

this model has been trained. So, we've

already fit it and now we're using it to

predict, right? And so, we're predicting

the sales and plotting that on the

y-axis.

So the sales are we're using the

predicted sales there which is our blue

line. So this is our line of best fit.

So this is our model prediction.

This is our model predictions. Right?

You can see it's a pretty decent uh

line, right? Pretty decent line of best

fit. Of course, there's some error here

like there, you know, it's not perfect,

but it it does a decent job of being a

best fit line.

Okay, so look how easy that was to

just to recap this to fit our model was

a linear regression.fit. And of course,

we're going to do more examples. So no

worries uh on that. We're going to see

this many many many times throughout

this notebook. But we have linear

regression.fit to train it. And then we

have linear regression.predict

to and we pass in our features and that

generates a predicted output.

Right. So what this is actually doing is

is computing this quantity.

We could do either.

We could do either. Um, so we could do,

so one thing we could do is plot uh, so

we could swap it out. We, we could do

either one. It doesn't, it's not a big

deal to do the training set. We could

do, so we could plot X test and then we

could plot linear regression X test.

So it's it's a similar line. Um it's

just different input features, but the

line is going to be the same. Just

different inputs,

but the coefficients are the same,

right? It's the same line. It's just we

generate different outputs.

So yeah, you could do either one.

This is this is honestly this is

probably better. I see what you're

saying. This is probably better because

this is the line of best fit through

this data. So that probably makes sense

to do to do predict on the test set.

Agreed on that. Probably makes about

most sense.

But you could do either one.

Yeah, I think that would be the most I

think that makes the most sense is for

it to be on the same one just to

validate. So like we could do we could

do training here and then train and

train just to see how that data lines

up.

Really, what we're trying to do is have

our scattered data and then our line of

best fit on the same plot. That's all

we're trying to do, right? So, yeah, I

think I think they should be the same.

I think that makes sense.

These values

or which values do you want to see?

Yeah, we could uh we could generate

those if we just do um let's go down

here. So the the line values

um are going to be uh the prediction.

So, um the the uh test

predictions

equals um

test predictions equals linear

regression.predict x test and then we

could uh we could print out our test

predictions.

Yeah. So, we can see what those actual

values are on our uh on the test set.

Yeah.

Um, we will do that. Yeah. So, you

thought we were checking how well our

data was trained. We will do that. Yes.

We haven't learned how to evaluate this

yet. We're going to talk about that

coming up next. Yeah. We will do that.

We just haven't learned how to do proper

evaluation

of a regression model.

But yeah, it's something we're going to

talk about for sure

and see how to do in our code.

Okay.

All right. Any other uh questions on

this example?

Again big takeaways

fit to train it and then predict to use

it

predict on the features to use the model

and make predictions with it.

So here is an example we we made all the

predictions. This these are all the

values that are on that line.

These are all our predictions and notice

they this is a truly regression right?

These are all floatingoint values. Um,

so this is definitely a regression,

right?

Okay.

Uh, that's a good question. Um,

I'm not sure if there is

If there's like a verbose

there's not really no there's not really

a verbose you can I mean you can look at

the source code if you really want to

see you can view the source code to see

um how it's done I can tell you I mean

so generally linear regression is done

in two ways either you use a formula um

to to solve the optimization problem of

minimizing like this this distance from

the points to to the line. Um,

or you use something called gradient

descent, which is how a lot of these

things do it is they iterate through a

bunch of different iterations where they

update these weights according to um a

certain uh basically a gradient of the

the error function. The error function

in this case is the is the squared

distance from the line to the uh to to

the points.

So uh we can compute the gradient of

that and do um gradient descent. So if

you really want to look into it, I would

do some research on like linear

regression gradient descent.

Okay, linear regression gradient descent

to see how that's uh how that's being

done. Yeah, it it's it's a pretty simple

procedure. Um, again, you have the the

notion is that you want to minimize

minimize the loss or the error. Uh, in

this case, the loss is the square

distance. So, it's like um there's like

a it's a formula. It's like a sum of a

square distance from your prediction

um or your label sorry to your model

which is the beta 0 um plus beta 1 x1

plus beta 2 x2

etc like your model and then squared. So

this squared this is the squared

distance here and you're minimizing this

guy which is like a calculus problem.

You you find you basically find the this

is this is I'm getting so far into the

weeds of this, but this is like a

parabola and you work your way No, no,

you're good. It's it's it's a good

question. Um you work your way down to

the minimum of it. Does that make sense?

Like you're working your way down here

and you do that through a descent

process, like a descent iteration.

Um

so

that's how these are found.

Um, but you don't see that happening in

the background. But if you look at the

source code, it I guarantee you it would

be it's either going to be this or

they're going to use the they're going

to use a a a matrix formula to basically

solve an equation um that involves this

basically the derivative of this set

equal to zero and you find the minimum.

Either way, you're finding the minimum

of this.

Okay. But yeah, I don't think Psycharn

has like a uh maybe there's some type of

verbose flag you can look for.

I don't think they have that though. Not

that I've seen.

All right.

So I have uh an important um concept to

talk about next which is going to be uh

called overfitting and underfitting

um which is a really important concept

that's related to the training and test

data we just split apart to do

evaluation.

And um essentially the the issue with

machine learning is that it's not

perfect and it can struggle in different

ways. And the two ways that it primarily

struggles is going to be overfitting and

underfitting. So overfitting is a

situation where the model basically

memorizes the training data so well that

it's it fails to generalize to new

examples. So what we see with

overfitting is this exact sign here

where we have really good performance on

the training data. So when so when we do

that train test split we see a really

good accuracy or really low error on the

training data but it does not perform

anywhere near that on that test data

split. So what that means is that the

model is overfitting to the training

data. it's basically memorizing it and

it's not able to generalize very well.

Now, why does that happen? It's usually

because the model is way too complex.

And that means generally you need to do

something to reduce the complexity.

Either you need to use a simpler model

or you need to use some type of

technique to mitigate overfitting. And

we're going to we're going to study some

of those techniques coming up in this

notebook. Uh we might not get to it

today, but we're going to study

particularly what can we do to prevent

overfitting because overfitting is the

more common issue with machine learning

models. They tend to do so well at

learning from data that they pick up on

small details and patterns in the

training examples that they're exposed

to. They don't do a great job at

generalizing to new examples. they can

struggle with that. So that's

overfitting is struggling to generalize

to new examples, but you do really well

on your training data. So it appears

like you have a good model, but it it's

not able to go and make predictions on

test data very well, which means we

would not want to use that model in the

real world, right? Because it's not able

to generalize outside of what it's

already seen. And that's not a good

thing if we're trying to use it for real

world examples, right?

So overfitting is a real issue. Um you

see it all the time. I've seen it many

many times in the real world, real

industry uh work that I've done.

Overfitting is a is a challenge for a

lot of machine learning models. And so

we need some techniques to overcome

overfitting and we're going to study

some of those uh coming up shortly.

Um, one of the things that we can do,

one of the one of the things that we can

do to detect overfitting is exactly what

we just did, which is you split apart

your data into training and testing so

that you have a chance to do an

evaluation to see if you're even

overfitting in the first place. You want

to see that performance be consistent

from train to test, right? You want to

see consistency. What you don't want to

see is performance that drops off on the

test data. It's much worse. You don't

want to see that. That means that your

model is overfit uh to your training

data and it's not going to perform well

in the real world.

Okay. So, we're going to have a couple

ways to uh overcome that. Talk about

that. Um now, the opposite can actually

happen as well, which is called

underfitting.

And underfitting

refers to the fact that a model is too

simple and it actually just performs

poorly across the board. So if we see

poor performance on the training and

testing data, that's a good signal that

the model's underfit and that means it's

too simple usually and you should try

using something more complex. Um, so the

best way to combat underfitting is to

use a more complex model. And as we go

through and learn about the models,

we're going to learn about which ones

are simple and which ones are complex.

So we're going to have a scale of kind

of complexity. And if you're

underfitting, you want to bump up to the

to a more complex model. If you're if

you're overfitting, one way of combating

that is to actually go down to something

more simple. Go the opposite way to

something simpler. So we need to learn

right now we've only learned linear

regression

but we will learn other models you know

in the future and we'll we'll talk about

uh their complexity and how they're

related to each other.

Okay, but these are two issues we see

just to draw that out again is if we

have a train test split where we have

7030 split let's say and we perform

really well over here but we go to apply

that model over here and it fails it's

accuracy drops off significantly more

error that's that's definitely

overfitting which is not good

right and then underfitting is just not

performing well in either case so even

on the training data itself self your

your accuracy is not very good. So

you're not really learning effectively.

You're underfitting your model. So

that's that's um underfitting case.

Okay.

All right. Now the issue is that it can

be very difficult to balance these two

and get it correct. That's what makes

machine learning a little bit

challenging is getting this balance

correct of simplicity and complexity. So

you don't want to be overly complex that

you overfit, but you don't want to be

overly simple that you underfit and

you're not able to learn effectively. So

there's a bit of a tradeoff there. And

this trade-off is typically known in the

community as bias variance trade-off. Um

in which case, uh it's basically like a

complexity simplicity trade-off. It's

another word for that. Um,

and so, uh, it's it's thought that, um,

if you, uh, if you have very, um, if you

have a situation where you're able to

fit the training data very well, you

risk not being able to generalize. In

other words, you risk overfitting, and

it's hard to um, it's hard to combat

that in a way. Um, and um, on the

reverse side, if you have something

really simple, um, you risk not learning

enough. Even if you're trying to combat

that overfitting, you risk not learning

enough and your model just doesn't

perform as well as it could. So, there's

a bit of a trade-off there of trying to

find the right balance between something

complex enough to learn, but something

not overly complex that it's going to

not generalize to new data. That's the

challenge. Um, like I said, we are going

to have techniques to overcome this. So

luckily there are things to basically

overcome this trade-off and um and help

us along the way so that we don't

overfit. They basically prevent

overfitting

um and allow us to use complex enough

models um that that won't be overfit.

This is in the um this was in our uh

lesson 3.2 notebook. So you want to pull

that one back up. We were working on

Monday.

Um, and just to recap this a little bit,

remember we were building a linear

regression, I wanted to recap some of

the steps we took there, um, that we

will be doing over and over again. And

really the same kind of steps, uh, that

we do here, we'll do in a lot of our

model building. Pretty much all of our

model building um, that we do, whether

it's regression or classification,

doesn't really matter. um we'll still be

doing a lot of these steps which are um

remember first we split apart our data

into kind of a features and a label

uh x and y and the reason that's

important is because um the model

training uses the features and the label

um to help train the model right they

use those separately um so we want to

split those apart whenever we can and so

we have usually Uh it's a good practice

to call your features capital X and your

labels lowercase Y. And what we do with

that is remember we immediately split

that into what we called a training and

a test set. And the picture we had for

that was something like this

where we had about 70% of the data

we used to train the model against and

then the other 30% of the data we use to

test the model against. Meaning that we

build a model over here and we apply it

to this set over here um to make

predictions. And then the that's where

the supervised learning really comes

into play, right? is on this test set.

We already have the answers. We already

have the label. And so we can apply our

model to this to the features over here.

Predict uh what the the label should be

and compare that. We can get a a metric,

right, that compares how close we are in

our prediction to the actual values. Um

and that was some of our performance

metrics. I'll recap some of those that

kind of measure that distance away from

our predictions to what the actual label

is. Um, but remember we had this train

test split function which helps us split

apart our features and our labels into

these uh four sets of data. So we have

our training features, our testing

features and then our training labels

and our testing labels. So we have all

of those and um really these two guys

are going to be used to train the model.

That's why they're called underscore

train. They're going to be used to train

that model and then the then we're going

to predict on these set of features and

then com use those predictions to

compare to this set of labels right

that's on the test test set. Um and you

notice here our test size is set to 30%.

Um, that's a pretty standard number.

Anywhere between like 20 to 30% is

pretty standard. Um, we'll typically

use.3, but it could be 02. Anywhere in

between is fine.

Okay, so we had that. Hopefully that uh

we remember that from Monday.

So we had a train and a test set. And

then building the model was actually

really really easy. Once you have those

train and test sets, um, we just import

our model object. So from uh scikitlearn

sklearn

um linear model uh module from that

package we import the linear regression

model and then we do um linear

regression.fit

and we pass in our features and our

labels and this is again this is where

that supervised learning is really

coming into play because we're passing

in these labels.

That's really what makes this work,

right? We need those labels to help

guide the model to make those updates.

If you guys remember, the model is

something that looks like this.

So, this was a bunch of different

coefficients

um times the features,

however many we have. Um, and so these

labels are really taking the place of

this and they're helping us um make the

correct updates to these to these

coefficients or sometimes we call them

weights. Um, these B 0, B1, B2. Um, we

find out what the optimal one is to get

the best fit, right? To get the line of

best fit. Um, that's what the model

training when we call this fit. That's

really what it's doing in the background

is finding all those coefficients,

right, to end up with the line of best

fit that has the lowest amount of error.

Okay, so hopefully that makes sense.

That's just a fit um to train our

models. And that's really going to be um

the case for

uh pretty much every single model that

we uh train with scikitlearn. It's

pretty much going to be a fit. we pass

in our training uh features and our

training labels.

Okay, so we had that and this was the

visualization of that where we had our

test points kind of scattered and we see

our line of best fit is the one that

goes through there with that minimal

error. That's that's the whole goal.

Pretty decent predictor.

Okay. And then we talked about

overfitting, underfitting. So just to

recap this, overfitting is the concept

of our model basically memorizing our

training data. It performs really well

on that training set, but it is not able

to generalize outside of that. So it

performs poorly on the test set or data

that it's never seen before. Um, and

that's overfitting. So the reason that

it overfits is generally the model is

too complex and it needs to be um it

needs to be simplified a bit. And one of

the things we're going to do today is

see a couple of ways we can alter the

linear regression model um if we are

overfitting to prevent overfitting. Um

so there's going to be ways to handle

this. Um and so we're going to explore

some of those today.

Uh underfitting is kind of the reverse

of that. Remember it's where the model

is not learning enough. So the

performance is poor even on the training

data. It's not good on the test data

either. Um that is a sign that the model

is probably too simple and maybe we

should use something more complex like

go from a linear regression maybe use a

polomial regression. Um or maybe use an

entirely different model altogether. Um,

if we're underfitting, our performance

is poor, it's a good signal we should

try something else. Um,

okay.

So, we talked about those

and one of the things we also talked

about was evaluations. If you guys

remember, we had different metrics that

we could compute to get a gauge of how

good our model is actually performing.

Um, one of those was MSE, which is this

mean squared error function. Um so we

did this example during class last time

on Monday um where we uh were able to

generate the mean squared error. That's

one of our metrics. And we can see what

the mean squared error is on the

training set and see what it is on the

test set by um just passing in our um

training predictions and our training

labels, our test predictions and our

test labels. pass those into this mean

squared error function and it computes

the MSE and that's that's a helpful

function from the scikitlearn metrics

um package um or module I should say and

we'll be using that quite a bit to do

you know evaluation of of especially of

regression right mean squared error is

pretty is probably the most common uh

performance metric we can have and if

you guys remember what it's really doing

is measuring these distances So mean

squared error is kind of like the

average distance away from our our

points to the actual um to the

predictions which the predictions are

all on this line. Um so it's like

measuring on average how how much error

do we have on average right? Um, and the

idea is the closer to zero the better.

Generally means that the distance away

from our prediction to our points is

pretty low. The closer to zero it is.

Um, which is pretty desirable.

So a low MSE is kind of what we're

looking for. Um, closer to zero the

better. And so um if one model has if

one model has um a low lower MSE than

another, it's it's a better performing

model, right? It has less error.

Okay. And then we also looked at the R R

squared or sometimes known as R2 um

score. Um this is another metric that we

could use that measures the the

variability

um of uh the predictions and if our

model is capturing that variability um

well um and so R squ is has a range of 0

to one one is better that means the

model is capturing the the changes in in

the um output it um our predictions

follow along with those same changes um

so they're pretty close um so closer to

one would be a better score. So we have

those kind of metrics. So like on this

data um this would this would show that

this model was underfitting remember

because this

mean this MSE was bad and this MSE was

bad.

Um and what we should think of these in

the units of what our labels are. um

especially if we take the square root of

this the RMSSE that was another metric

we had um the square root of this is

actually in the exact units that we um

have for our labels. So uh in this

example this was the um this was the the

units or the sales versus the TV

products, right? Um and so this would

indicate that on average if we take the

square root of this um

in the square root of this um we have uh

um we're on average about 11 sales units

off squared. So if we take the square

roo of that um it's somewhere around 3

to four um somewhere in between three

and four units off. And this is as well.

Um, and because both of these are still

not close to zero, um, this would be

under fit. And this shows that as well.

This isn't that close to one. It's

decent, but it's not, um, not that close

to one. So, we would say, and

performance is poor on both training and

test sets. That's the key indicator of

underfitting. It's poor on both.

Yeah, exactly. High MSE correlates to

underfitting. Yes. Yes. And it what's

key is it's high MSE on both on both the

training and the test sets.

If you have a high MSE on your test set

but a low MSE on your training set,

that's overfitting, right? Where it's

not generalizing from the training set

to the test data that it hasn't seen

before. That's overfitting. So the key

is high MSE on both sets.

All right. So we talked about that. Um

we did polomial regression last time. So

that was um doing

that was uh making a curved graph um by

transforming the features into polomial

features and then doing linear

regression with that. So you guys

remember from Monday we did this where

um we took our features and uh

transformed them according to this

polomial features from scikitlearn. So

we can go all the way up to degree

whatever degree we want. So we put in

four here but there's nothing special

about four really. This is just testing

it out. um and we generate the the

polomial features and we can fit a

linear regression on those polomial

features and we get a slightly better

model, right? Um it fits the data a

little bit better than just a straight

line. This curved line with the polomial

features um performs a little bit better

and we could see that with the MSE,

right? or we could evaluate the MSE of

this um and it would be lower.

It would be lower than the curve line.

And so that's something we could do. Um

we would just have to pass in these test

predictions, the training predictions

and then the the test labels and

training labels and pass those into the

mean squared error function and we could

compute that, right? Wouldn't be hard to

do.

All right. And then finally where we

left off um you know is on our

performance metrics. So we talked about

mean squared error. That's that average

distance away from the labels to our

predictions. Um and we take the square

root of that. It's it's basically

measuring the same thing but it's the

square root of it is um more

interpretable because it's in the same

units as our label.

um mean absolute error is is the average

distance of the absolute value. So it's

not the squared distance formula like a

uklidian distance but it is a absolute

value. So it's a little bit um less

sensitive to outliers. They don't get

magnified as much. Um but it's not

typically used as much as a mean squared

error would be with regression. um we

talked about the last time because um

the distance formula or that distance is

actually what's used to train the model.

So it's a more natural um fit for a

performance metric for it.

All right. And then we had R square. We

just talked about that closer to zero

would be um worse. Closer to one would

be better. That means that the model

explains um all the variability in the

in the predictions. Uh it captures those

predictions um closely to the labels

um very well. So uh one would be better.

Closer to one would be better.

All right. So that's where we left off.

Um we're gonna pick up from there with

cross validation. um we've actually

already seen one method of cross

validation. So we're going to study um

we're going to kind of recap that and

and then um talk about cross validation

in general um and look at some more

sophisticated techniques of it um coming

up next. But before I do that, any

questions about anything we've covered

um to this point in in the recap or

anything from Monday? Any questions on

that?

All right. So let's talk about uh cross

validation. Um now this term cross

validation refers to a technique that

evaluates performance. And what it does

is it divides our data into essentially

um training and test sets which we've

kind of already seen. And then we are

able to train a model on on the training

set, evaluate it on the test set. And

that's where that's where we get the

name cross validation because we're

crossing over our model from one batch

of data used to train it over to another

set of data used to validate those

predictions. Um, and there's actually

different ways to do cross validation.

So cross validation is a bit of an

umbrella term for multiple ways to do

that. We've already seen one way of

doing that um which I'm going to scroll

down to is um known as a hold out cross

validation. So that's um what we've been

doing so far. So this is just um

generating a train and a test set

train um split.

Um that's the that's what's known as the

hold out cross validation method. Um and

and this is exactly what we've been

doing so far, which is you split your

data into some type of split, usually

7030,

um of a train and test

and then you um train your model on this

section of data and then apply it to

this to evaluate performance. Right? So

that's that's what's known as the hold

out method. Um it is uh you know

relatively simple. It's pretty fast to

do. Um, but there are more robust ways

to try to divide up our data a little

bit uh more evenly. Instead of just

having one split, we can actually do

many splits, which is the idea of um the

next kind of cross validation I'll

cover. But hold out method is one that

we've already studied. It's the most

basic type of cross validation you can

have. Um so hold out this is the most

basic

and we we've already been we've already

been uh working with this type. Okay.

So we've we've already seen hold out

method. Let me uh explain to you a more

sophisticated method a little bit more

advanced of a cross validation um which

is known as Kfold cross validation. So

this is um going to be a little bit more

advanced of a technique but this is the

idea of kfold is that you take your data

set

and you split it into k number of what

are called splits or folds. So you take

your data and you let's say it was let's

say k equals 5. So we have five splits

here.

Okay. So let's say k equals 5. We have

five splits. So what we're going to do

is we're going to we're going to train

our model on K minus one of those folds.

So if K was five, we had five splits.

We're going to take our model and train

it on four out of five of those uh

splits. So let's say it's these four.

We'll train it on these four.

Okay. And then what we do is the one

split that's left over, we will we will

test our model against that split. So

we'll test here.

Okay. Now, this sounds very similar to

the hold out method where we're doing a

train test split, but it's a little bit

this kful cross validation a little bit

more sophisticated because we repeat

this process that I just mentioned over

and over for all combinations of the

splits. So then what we'll do, this is

just one trial that we'll do it again,

but this time we will pick um four

different splits. So, this time we might

pick,

let me do blue. This time we might pick

this one, this one,

um,

this one,

and this one.

And then those four we will train our

data on. And then we will test against

this one. Okay. And we'll do we'll

repeat this

repeat for all combos of the folds.

Okay. So we'll repeat that. So

essentially what we're doing is rotating

through. Every time we rotate through

one of the folds is going to be left out

as a test set. Now this is a little bit

more robust than just a train test

split, right? because we are exposing

our model to more of the data in in

doing this, right? Because we're going

to split it evenly into five or 10

splits. Those are pretty common um

number of folds to use. 10 or five. Um

those are the ones I've most commonly

seen. Um but we're going to by rotating

through which folds are being used for

training, which ones being left out. um

we are exposing our our model to more of

the data this way than just doing a

single train test split. Right? So now

what do we do with with the results is

every time we do this we we generate um

an MSE let's say or some type of

performance metric. So let's say we

generate an MSE from this guy

we generate an MSE from this version and

we generate an MSE for all combos.

each combo we generate MSE and then what

we do is we average

the metrics

or the in this case uh if we use MSE we

would average those together. So every

time we do a fold combination and we

keep four of them for training, one for

test and we rotate through all those

combinations, we are going to generate

an MSE for every combination

then we're just going to average those

MSSE's to get a final. So the final MSE

of cross val of this K-fold.

So the final metric

is just the average of the uh

performance on all of the fold

combinations. Okay. So our final MSE, we

just average all those MSE from all of

our combinations.

Okay.

Now, what's the advantage to doing this?

It's way more robust of a estimate of

the of the performance of the model

because we're exposing it to all

basically all of our data, right? We're

getting a sense of how it performs

across all those different folds. Um

rather than just doing a single train

test split, which is a bit it's basic,

it works, but it's a bit basic. Um so

this is more robust estimate of the

performance.

Now, what's the drawback to doing this

is that it's more intensive. So, if you

have a lot of data, this is going to be

pretty expensive to do because you're

going to have to especially you have a

high number of folds, right? You're

going to have to divide your data into k

number of folds and you're going to have

to do this over and over again. Um, and

if it's a large data set, it might take

your model a long time to train. It's

going to be a little bit more uh

computationally intense than if we just

did a train test split.

Okay, we just did a single like 7030

split. We only do that once. We only

train the model once, right? We train it

on the 70, apply it to the 30% test data

and evaluate performance that way. Um,

so we're only really using the model and

training the model once, but in this

kfold, we're going to do it um, you

know, k number of times essentially

or I should say one for every

combination that we have to work through

of of all the folds.

Okay.

All right. Does that make sense? Any any

questions on kf fold cross validation?

So k K is an important uh number here.

It it's how many folds how many splits

do you have? A typical value for K is

going to be somewhere like five or 10.

So 10 folds or five folds. Those are

pretty pretty standard

from what from what I've seen.

But does the does the concept make sense

or is there any questions on it on in

terms of um you're always going to leave

one fold out. You're going to split it

up into K number of folds. Always leave

one out. Train on the rest of it.

Evaluate on that one that gets left out

and then rotate those through. And

you're going to do that for every

combination and average all those

metrics.

And by the way, there's going to be an

easy function in scikitlearn that will

do this for us. So managing all these

combinations will be really easy. It's

actually just built into scikitlearn. So

we don't have to um we don't have to do

this all by hand. Okay, this will be in

scikitlearn. It'll handle doing all

these combinations of folds for us and

computing the average metric will be

really easy. So um

we don't have to worry about that. We're

going to see an example of this coming

up shortly.

All right, of kfold cross validation,

but this is a this is a really widely

used technique. And again, like the

purpose, you may be wondering like

what's the purpose ultimately of doing

this? It's to get a sense of if our

model is going to perform well on new

data. That's really what we want to

know. Like is the model going to perform

well when I start to use it on new data

that it's never seen before? And this

kffold is a decent indicator of that

because we are varying which data it

sees across many different folds. Right?

So it's a it's kind of a good um proxy

to exposing it to different kinds of

data each time and seeing how it

performs.

All right? Because we're working our way

through each one of the folds. There's

always going to be one fold left out.

We're going to change which fold gets

left out each time. And um that's sort

of mimicking the idea of we're going to

apply our model to new data and see how

it performs. And it's it's new data

every fold.

um how we know which model is best suits

for which scenario because we have Yeah,

that's a good question. Um,

so my we're going to learn this as we go

along because we haven't covered all the

models yet, but generally the best

advice I can give on that is

you you generally want to start as

simple as you can get and then if it's

not performing well then work your way

up to something more complex.

So we are going to have models that are

simpler. We're going to have models that

are more complex. The rule of thumb is

to start with the most simple model that

works.

So you're usually going to have the same

ones that you're going to try in the

beginning. And linear regression is a

very simple model. It's usually the

first one you want to try for regression

because it's the simplest.

Um, and for classification, we're going

to have a similar like logistic

regression is the simplest kind of

classification model we could have. So

usually want to start with that and then

if it underfits like if we see it's

producing a lot of error then we work

our way up to a more sophisticated

model.

So um that's the way we that's the way

it should usually go is simple to

complex it based on their performance.

So we evaluate it and then we can repeat

the process. If it's not performing well

we can try something different that's

more complex if it's underfitting.

Uh this is a good question. Does a model

reset after training each K minus one

fold? Um yeah, it's essentially like a

blank model every time uh every fold. So

um we imagine like you have a brand you

have a fresh model every um k minus one

combination. Yes.

And the reason the reason it has to be

that way is because you don't want the

other folds influencing the model that

like on on the next combination. You

don't want the previous combination to

influence the results on the next one,

right? Um you want it to be a fresh

evaluation on every combination of

folds.

Okay.

All right. So, let me describe to you a

variation on what we just um talked

about with the K-fold. So, there's

another cross validation known as

stratified K-fold. And um this is the

same exact procedure as k-fold except

that when we this is used for

classification.

Um so when we do classification

uh we want to make sure that the

different categories are going to be um

split amongst those folds in a

proportional way. So we don't what we

don't want to happen is um when we split

apart the data. So, let's say we have

let's say we're predicting um spam not

spam. What we don't want to have happen

when we do our splits is we don't want

to have all of the spams end up in one

fold and then every other fold has no

spam, no spam, no spam, no spam, right?

That's not very good. Um because if we

if we train against all these guys, we

have no shot at predicting spam when

they've never seen spam before. So

stratify kffold is is used in

classification

and it's to um it's to make our splits

ensure that they have basically a

balanced number of categories for each

split. Um so that we don't end up with

certain splits with way more spams than

not spams. Um so we we do what's called

stratifying where we make sure the

proportions are balanced across each uh

split. So this is only really useful in

classification, not really necessary in

regression because we're predicting a

value. But if we were predicting a

category,

like in classification like fraud, not

fraud, we don't want to do the split and

have every single fraud example um by

bad luck in our shuffling and split end

up in one split and every other um every

other split has no examples of fraud.

Right? Right. So we want to stratify

this to spread out those um frauds

against all the other splits. Um so uh

again um scikitlearn will take care of

that for you. Um but if you're doing

classification and you have an

imbalanced data set um you you really

want to make sure you stratify kfold. um

imbalanced meaning that you have a a um

different number. Like if you're doing

fraud, not fraud, you have way more not

frauds than frauds. Um when where that

category is imbalanced,

you want to make sure it's balanced

across all your splits.

Um so this is this is useful in

classification only, not really

regression, which is what we're talking

about right now. Um but it's just a

variation on this that ensures when we

do those folds um the data is

distributed evenly amongst those folds

as much as we can. The labels are I

should say.

Okay. So that's stratified kfold. It's

the same same procedure once we have our

splits. It's the same where we do k

minus one of them. We train test on that

last fold um and then rotate through all

the folds and and average all the

metrics. the same exact procedure. It's

just the splitting itself um is going to

be balanced in a stratified kfold.

Okay, so hold out we've already talked

about um is just doing a single train

test split. We've talked about that. One

more variation that is a bit of an

extreme version of Kfold. So it's

actually the same process as Kfold, but

it's an extreme version is if you set K

equal to the number of data points. So

you basically are um this is a really

really extreme kfold where you um

basically are training on all the data.

Um so you're training on all the data

except one point and then you test

against that one point. Um now why would

you ever do this? Um it's mainly so for

this reason here. it's to um maximize

the amount of training data that your

model gets exposed to because instead of

just doing instead of just doing five

splits

um which would be like

you know these four folds are going to

be used and then we um test against one

fold um we're essentially going to use

99% of the data right one point is going

to be left out 99% of the data gets used

to train um and then we're always going

to leave out one point and and the issue

is we're actually going to do that over

and over and over again and rotate that

one point to cover the whole data set.

So we're going to train on 99% leave one

that one point out

and then rotate through every

combination of points until we've left

out every single point and then average

all those together. Um so this is a this

is an extreme kffold. Again the number

of folds is actually equal to the number

of data points in this case. So we have

every point is its own fold and we train

on everything but one test on that one.

This gets you the maximum size of your

training data because you're basically

going to have every point but one used

in the training.

This gets you the maximum size. However,

it gets you the maximum uh expense

especially for large data sets. This is

going to be usually you're not going to

use this um especially for large data

sets because it's just too extreme. It's

going to take you a really long time to

work through every single point being

left out. Um it's just going to take a

while to do.

So for that reason, the leave one out um

that that's why it's called leave one

out because it's you're leaving one out

every single time. Um is rarely used. I

I don't really see it used that often,

but it is an extreme version of K-fold

cross validation.

Okay. But rarely ever actually used. I

think the the ones that get used the

most are definitely the hold out method

with just a regular train test split. Um

and then uh the other one that gets used

quite a bit is is Kfold

or stratified K-fold if you're if you're

doing classification, but certainly

K-fold in the in a regression case.

Okay.

All right. Um, we're going to do an

example with these guys. So, we'll do

that next. Um, with with the different

cross validation techniques. Um, but any

questions on what they are doing

conceptually before we actually do the

code example?

Okay,

very good.

All right, so let's see some examples.

Um let's go into our code and build a

model and do the different cross

validation techniques on it. Um you're

going to see it's actually going to be

really easy to do and we it sounds

complex like doing the kfold and leaving

one out and testing it sounds kind of

complex but I promise you scikitlearn

makes it really easy to do. Um

and so uh we won't need to do too much

besides just use the right uh tools from

scikitlearn. Uh so we're going to we're

going to see that. Um so here we have

some imports. The um primary uh thing

that's a little bit new for us is going

to be these um different kinds of cross

validation techniques. So we have our

kfold, we have our stratified kfold,

leave one out um which are those

different cross validation techniques.

Um these are going to be used in

combination with this cross val score

which is going to keep track of the

different um metrics and then average

them

uh while we do one of these um cross

validation techniques. So this guy gets

used in combination with one of these to

um as as we're going to see in the code

uh to average those metrics. um doing

the different folds, right? Perform

doing performance against the different

folds. Okay. And then of course we need

a model

using linear regression. That's that's

the one we've studied so far. Um and

then we have just a regular metrics if

we want to compute those. Um using maybe

just hold out, right? And hold out um

which which is just a regular train test

split. Um we could use these guys to

evaluate performance.

But in a more sophisticated K-fold style

of cross audition, we're going to use

this to evaluate the the performance.

Okay, let's see.

So, we're going to be working with this

housing with ocean proximity data. Um,

you guys should have this one. Uh, so

you guys should have this one. So, if

you want to follow along and run it

yourself, um, you can load that one in.

Um, I want to make sure that I have it.

Let me pull that one in. So, it should

be this guy.

I'm going to load that in so I can make

sure I run it with you guys.

Um,

so let me run this.

Do you guys have that data?

The housing with ocean proximity?

It's another it's another housing data

set. Um,

but it it's a little bit different than

the ones we've seen before. It has a a

special feature for how close it is to

the ocean, the different locations.

So, it looks kind of like this. If we

load it in and do our head, which is

usually what we do, right, we can see um

we can see that it's got these features.

So, it's got uh uh bedrooms, total

rooms,

um it's got uh median age. Now, this is

this is looks a little strange for total

rooms and um uh bedrooms and population,

etc., but it's um

it's it's got those uh it's got those

because it's representing an entire

neighborhood. So, it's an entire

neighborhood and we're looking at this

um this is actually going to be our

label is this median house value for the

entire neighborhood. So, what's that

median value uh in the neighborhood? And

this is the total number of bedrooms,

total number of rooms, um population,

households. So, how many houses are

there? Um median income. And of course,

these are scaled. So these are um likely

times you know uh thousands

um

but um that's our data. We could

describe it.

So we can see the average age median age

um which sounds a little um weird but

that's it's because again this is the

median of data within a neighborhood. Um

so the average of those is about 28 or

29. Um we have

u

total bedrooms. The we can look at the

min. There's some data that only has

one. So it's likely only one house in

there. Um which is what this represents.

There's only one house. So there there

is some neighborhood that only has one

house. Um, and we see the median, um, we

see the minimum, uh, median house values

there. And then the maximum down here,

um, is a pretty big number.

6,000 households is the largest that we

have in any any one of these

neighborhoods.

Okay. So, just a little bit of

description of the data.

Okay. So then we can run.info. So this

is um let me ask you guys, were you able

to load this? Were you able to run this?

If you're following along, were you able

to

load it and take a look at dot head.

Okay, great. Great.

Okay, so we're able to load that and

then look at dot head. Perfect. Um

okay.

Um and then we run describe which gives

us that uh usual kind of statistical

description. Uh so we can see some

interesting stats about those.

What do you guys notice about the info?

Anything interesting that we see from

there?

Is there any missing data?

Any features that have missing data? Can

we see

object? Yeah, object type usually is

string. If it's an object type, that

usually means string. Python when we

read it into pandas it usually is just a

string.

So that that makes sense like we have

mostly numerical features but then we

have a this ocean proximity which is a

string.

Yeah. Total bedrooms has nles. That's

right. Because you can see here this

does not equal the number of uh rows

that we have. So, this is the number of

rows. There's about 20,000 rows. That's

a good size data set, right? 20,000

rows. That's decent. Um, we're

definitely missing some data here for

sure. Um, we could count how much we're

missing exactly by running this is NATO

sum. Um,

and so we see that total bedrooms is

missing about 200 uh 200 rows are

missing total bedroom uh value.

Okay. And then one thing I wanted to

look at is yes, this is a string. So

what remember what we can do with those?

That's a categorical.

So ocean proximity

is a categorical

string

feature.

So we can take a look at its value

counts, which is usually a good idea to

take a look and see what possible values

that feature could be. So if we look at

our

um what are we calling this? Housing

data.

Housing data

ocean

proximity

dot value counts.

So here's the different types that that

one can be. So there's some

neighborhoods that are less than 1 hour

from the ocean. There's some that are

inland. There's some that are near the

ocean. There's some that are near a bay.

There's even five of them that are on an

island. So, these are the different

values of the ocean proximity. So,

remember, you can always do that. If you

see a string feature, you can always

take a look at what its um categories

are. And it looks like most things are

less than 1 hour from the ocean, but

it's kind of evenly distributed here. Um

otherwise

very few islands.

But as you can imagine like this feature

is probably going to be important for

determining um what the value is, right?

Probably going to be important.

Okay. So, um, we need to deal with these

NLES. If we're going to build a model,

right? So, um, this is all of our

typical data prep. If we want to build a

model, we're going to have to deal with

these NLES. What do you guys think we

should do with the NLES? What would you

what do you think for total bedrooms?

What do you think is a good strategy to

do? Keep in mind, we have 20,000 points,

20,000 rows I should say, and about 200

of them are null.

Right. So about 200 are null. Um so what

do you what do you guys think would be

like a good strategy to deal with those

nles in that case?

average. We can't ignore it because we

can't ignore that column.

We can't ignore the whole column. So,

something needs to go there.

Probably don't want to make it zero.

I think average is a decent average is a

decent idea. Probably don't want to make

it zero because um that would indicate

that there's no bedrooms and yet we

still have a bunch of total rooms. So,

it probably doesn't make sense to do

zero.

Average, I think average could be a

decent one.

Now, in this example, what we're

actually going to do is we're

rows.

We're actually going to drop the rows al

together. Now, why are we doing that?

It's because we have so much data and

only 200 of them are null.

Okay, only 200 of them are null. So,

we're actually just going to drop the

rows. Now, that's a choice.

Um, that's a choice, right? Is that we

could fill in with the average like you

guys are suggesting. What we're actually

going to do is just drop the rows. It It

makes up less. It makes up about 1% of

the whole data. So it's not that much of

it is missing. We can drop those rows.

So that's actually what we're going to

do here is we remove all the roles with

the NLES by doing drop NA. So this just

drops them. So those rows are cut out.

Um, it's arguable that we could replace

it's arguable that we could just replace

it with something and I think you guys

have good thoughts which is the average

a default

um assume total bedrooms. We could we

could try that. Yeah.

Assign a value based on comparable home

value. Yes, you could do that too.

That's a good strategy is to look at the

other rows that are similar to it and

fill in a value. That's absolutely fair.

Um, in this example, we're actually just

going to drop those rows,

but I think that's totally um totally

valid.

This is a choice.

We could fill NA with different values

such as the average,

total bedrooms,

um, derive a value, etc. So, we could

derive something, which I think Brent,

you have a good suggestion. That's a

good suggestion. Um, we could derive

something like that, uh, and fill in the

blank, and that's I think that's totally

valid. Um, we could take the average of

the um bedrooms. Uh, I meant total rooms

here. Sorry, total rooms. Um, we could

fill in we could fill it in with the

total rooms for that category um or for

that row. Um, many options. In this

case, we're actually just going to drop

those rows because they make up such a

small percentage relative to the 20,000

rows that we have. It's about 1%, right?

200 rows is about 1% of 20,000.

So, we're just going to drop them. But

that's a choice. We don't have to drop

them. We could fill in with something.

Um, and if we did that, we would use

fill na rather than drop NA, right?

Uh after dropping the rows, how many? So

it's just so after we drop the rows, um

after we drop the rows, it's just going

to be we still have all our other rows

are intact, right? So if we look at this

now,

we now have um slightly uh slightly less

entries.

So now we have this this many um rather

than rather than this many,

right? We dropped those 200

But they're all filled in. Yeah, they're

So all the other columns are still

filled in. We're just we're we're

cutting out the whole row. So if you

think about our data set, um we have all

these rows and all these columns. What

we're doing is like if there's a null

here, we're just we're just getting rid

of that whole row, right? And so we

still have all the other rows intact.

Uh, we can drop them because we have a

good sample size. Yes,

that's exactly right, Ronald. Yep, we

can drop them because we have we have

20,000 rows and only 200 are missing

values. So, that's totally fine.

Uh, drop a removes all rows that has any

null. Yes, that's true. It it will go

ahead and just drop any row where

there's any null, no matter what column

it's in. Yes,

index. Yeah, the index is not getting

reset. Um that's true. So um what we

what you can always do is you can reset

the index. So, um, if you want to, it's

optional. We we're not really going to

use the index for anything that

important, right? But what we could do

is, uh, reset index.

Uh,

we could do that, right? Which will

reset it.

So now now it gets reset.

But um let me actually I don't I don't

really want to do that. I'm going to

reset this.

Um,

yeah, we could do that.

Okay.

So now importantly there should be uh no

missing data of this of this new one

where we've dropped NAS right. So now

this is good. If you now the reason we

had to do this is because if we try to

build a linear regression and we have

nles in there um the the issue is like

how do you build a model where you have

something like this

and these are null like what do how do

you multiply a number by a null?

Um, we can't really do that, right?

We can't really do that. So, um,

so therefore, uh, we need to get rid of

NLES like the the null's not really

going to work in there. So, uh, we need

to get rid of them for linear regression

to to really have a chance to work,

right? To train it and be able to use

it.

You got to get rid of those nles.

All right,

any questions so far? So, we haven't

done any modeling yet. We're doing some

We're doing some data preparation before

we get to the modeling. And we haven't

done any cross validation yet. We

haven't set that up. We're just doing

our data preparation before we get to

the modeling. Right. So, we've dropped

some NAS. We've checked it. Um, we're

going to do one more prep step, which is

to um change that ocean proximity

feature into something numerical because

again, how do you build a model where

you're inserting a string into those

like beta 1, beta 2, beta 3 times of

features? You can't really do that when

it's a string. Um, so what we're going

to do, and I'm going to get rid of this

because I don't think we really need

that. um is we are going to uh run this

get dummies function which is our um our

get dummies function is our usual one to

uh our git dummies one is our usual one

to um

uh get our one hot encoding. So this is

our uh one hot encoding here.

We now are going to have data that's

like this, right? So we have ocean. So

So by the way, this prefix

um this prefix is OP, which which is

short for ocean proximity, right? So we

have ocean proximity uh less than 1 hour

from the ocean, ocean proximity inland,

ocean proximity island, near bay, near

ocean. So these first five rows are near

the bay. Um so they have a one there and

a zero in the other spots. So this is

good. This one hot encodes that feature

into these numerical uh values,

right?

Were you guys able to run that one? The

get dummies

So the reason that Yeah, that's a great

question. How did it go ocean proximity?

It's because um that is the only uh

string feature we have. That's the only

one we have. So it it's going to look

for any non-numericals and one hot

encode those however many however many

there are. So whatever objects we have

which are strings, it's going to

automatically oneh hot encode those.

Yeah, we could have right we could have

went here and did Right. We could have

done ocean

proximity,

but we only have one of those features.

So, it's just going to do that to the

whole data frame

uh on that one feature. So, what we're

going to do is um go ahead and split it

into an X and a Y um which the X is

always what includes our features. The Y

is what we are trying to predict, which

is the label. Now, um, in order to

separate those out, what we're going to

do is assign X to be the variable that

is, um, our data frame minus this median

house value column. So what this is

doing is um uh it's not permanently

dropping because we're not uh dropping

it in place but it is returning us a

copy of the data frame with the median

house value column left out right it's

dropped. So this is this is uh something

we want to do because that will the rest

of it will contain our features, right?

So um this will temporarily or I should

say return a copy of the DF with um

median house value

dropped,

right? Median house value dropped. Um so

we go ahead and drop that one. Uh now

remember it's not permanent. It's just

giving us uh the remainder of it which

is this housing data dropping this and

it's assigning that to x and then we're

taking the actual median house value

column from the original data and

assigning that to y. So this is going to

be our labels

right. So this is what we are trying to

predict.

Okay, so that is our Y and that's always

how it is. X is our features, Y is our

labels. Um, hopefully that makes sense.

What this is doing is this is going to

get rid of that label column and

everything else will be our features.

And then this will get rid of this will

just assign the label column to Y.

All right. And then what we can do is

pass X and Y into our train test split

function. And this will generate the

hold out set. So if we want to do the

hold out cross validation, this is how

we would do it is we would split the

data into X-ray, X test, Y train, Y test

um using train test split. So this is

what we did last time. This would be

this would be for hold out cross

validation,

right? where we are uh uh just have that

one one set for testing and one set for

uh one set for training one test one set

for testing I should say right so this

is pretty standard train test split um

we pass in that x we pass in the y we

use a 30% test size is pretty standard

and random state so that we get the

consistent shuffling if we were to run

this multiple times um we we get that uh

consistent randomization

Okay,

so we have that and so now our X train

is a percentage um of the data frame of

the 20,000 uh rows and the X test is uh

30% of that. So it's only about 6,000

rows which is what um the shape of that

is.

Yeah. X so X is our features. So we're

we're putting all of our data in that is

our features into X. And so the the um

most efficient way of doing that is um

the most efficient way of doing that is

to

uh just take our data and drop the

median house value column because that's

our label column. So we just remove

that. The rest of the data is our

features. So that's what that's what

this X is, right? It's all of our

feature data. All of our columns that is

not the label column essentially is what

that's doing. And then Y is our label

column from our original data.

Right? Y is our label column. And so

this this will um contain all of our

labels which is the median house value.

X X X contains every column but the one

we're gonna so we we ultimately decide

that but X contains um X is everything

that is not our dependent variable which

is what we're predicting. So we're

removing what we are trying to predict

from X. X should be everything else.

That's always how it's going to be. X is

X is always going to be all of those

independent variables that we're using

to predict the median house value. So we

are going to predict the median house

value. We need to remove it from X.

So we're we're taking everything but

that column.

So it's the whole data frame. It's the

whole data frame minus this one column

with just the dependent variable. Right.

Exactly right. Removing the dependent

variable and keeping all the

independence. That's exactly right.

Exactly right. So think about it in

terms of the model. Let's go back to the

features. Right. Think about it in terms

of the model. We are trying to predict

this this value. We're building a model

to try to predict this. So we are going

to make sure X is everything but this

right. So this is actually just Y.

That's our label. That's our dependent

variable. Right? That's Y. Everything

else is belongs to X. Everything else

belongs to X including all of these.

Right? We choose this one to be Y

because we're building a model to

predict that. That's our label.

All right. So, we have our we use X and

Y to do our train test split. So, we

have our our training features and our

test features and then our training

label and test labels here. Um, pretty

standard there.

Um, okay. So, this is what's new is if

we want to do K-fold uh validation, what

we're going to do is create a kfold

object. So, we have this kffold from

scikitlearn that we already imported. we

are going to create a kfold um where we

are going to specify how many folds we

want. So that is the in uh inslits

parameter as this says um this is going

to be uh uh in this case we're going to

do 10 folds. That's pretty standard. So

I think the typical number of folds that

I've seen and I've worked with in my in

my career is usually five or 10.

Five or 10 folds is the standard.

Okay. So, we're doing 10 folds in this

case and we're setting a random state

because we're going to do shuffling. So,

in order to produce those folds, we're

going to shuffle the data first and then

split it into five folds, right? So,

this this kfold object is going to

manage creating these splits for us,

right? These even splits. I know I I

didn't draw it even, but um it's going

to manage these five folds for us and

it's going to shuffle the data and

assign them to these different folds and

we're and then what we're going to do is

use those to do our training. We're

going to execute the cross validation

using this k-fold object.

Okay, so we create the kffold

um we initialize our model as well. So,

of course, in order to train something

uh in the Kfolds, we're going to need a

model. In this case, we're using linear

regression, right? Which is which is the

model we've been studying so far. So,

you have a linear regression. Um now,

look how easy it's going to be in order

to execute cross validation. All we need

to do is um all we need to do is create

a crossfile score function.

um or I should say use the cross val

score function from scikitlearn. So we

use that with the model we want to

train. So our model goes first. So

that's the linear regression object.

Then our data. So our extra our features

and our label for our training.

And then um let me skip over this for a

second. I'll explain what this is in a

second. Um but then we are using uh the

cross validation technique is our kfold.

So this is where our k-fold object goes

in the CV parameter which is cross

validation. So what cross validation

strategy are you using? We're using

kfold and the kfold we're using is this

one we defined up here KF. So we're

putting that right here for this. And

then um in jobs um allows us to

parallelize this. So if we set it to

negative one, that's the that that's the

default. Um it will do it will actually

train across the different combinations

in parallel. Um which speeds it up. So

you want to you want to keep this to

negative one if you can. So um now let

me describe the scoring. So what this

means is we put in our metric here. Um

and so you can put mean absolute error,

you can put in mean squared error. Um

those are the two that we can use. And

um the reason we it has a negative in

front of it is because we want to find

the one that has the lowest score.

That's going to be our best model is the

one that has the lowest score. So we

take the absolute value.

I'm sorry. we take the abs the the the

metric and we take the negative of it um

because the highest scoring one is going

to be the closest to zero. Um so it's

just a we use the we use the negative of

the of the metric um because on the

number line like the the highest um

scoring one should be the least um or I

should say the maximum negative that we

can get. that's going to be closest to

this to zero. So if here's zero, this

would be like -1 is better than -10,

right? So something that scores um the

maximum negative uh absolute error would

be closest to zero

and something that has more is going to

be on this side.

So this is only the reason we need this

is only just to keep track of the scores

of each individual um fold. Okay.

So the one so the reason we can do that

is at the end we can kind of see which

which combination performed the best. Um

it's going to be the one that has the

highest uh highest value of the negative

which is closest to zero.

That's just a convention.

Yeah, it's just because um it's because

the cross validation is looking to

maximize the metric. So, whatever has

the best score

um whatever has the best score is

considered the best uh performance. Um

but we are using uh something where

lower is better. So we we take the

negative and like the the highest

negative would be closest to zero,

right? The highest negative is going to

be closest to zero.

So that so it's it's just because like

we want the lower score to be the best.

The lowest score should be the best.

So we take the negative of it. Um and so

something that is more negative is going

to be worse. Yeah, that's the reason.

So something that's down this way is

going to be worse. Okay, so it runs this

and what you can see is if we actually

print this out, if we print out our

k-fold scores, what we should get is 10

different scores.

And you can see um we have 10 different

uh scores here which are all negative

because we're taking the negative of the

absolute of the mean absolute error. Um

so what we would be looking for here is

um we want to take the average of these

scores but take the absolute value of

them to get the best performance. So

this is capturing like this is the score

on the first fold combination. This is

the score on the second fold

combination. This is the score on the

third fold combination and on and on and

on. And these are the absolute errors.

Okay, these are the absolute errors. Um

so if we take a look at computing the uh

average, which by the way, we don't need

this import because we're using the

numpy average. So that's fine. um we can

take the absolute value of those um and

take a look at the average MSE

or sorry MAE. Now I want you to think

about this this uh average performance.

So this is our performance right here on

the cross validation.

This is our average

MAE

across all of our fold combinations. So

that's a that's an indicator of our

performance, right? um for the cross

validation.

Now, what are the units of our original

uh the original median value? They're

already in the thousands, right? So, if

we go to that feature, they're already

in these hundreds of thousands. So, this

is not a very good error. It's it's kind

of high, right? because it's in this is

49,000.

Um that's that's how far away we are in

absolute value on average is $49,000 um

dollars on the median value. That's not

very good. So this score

this score is

um not very good. So this model is not

performing that well and we can see that

by comparing this error to our actual uh

data. So this is right around 50,000

and our median uh house values are in

the hundreds of thousands. So on average

we're 50,000 off when we make a

prediction. That's a significant amount,

right? It's a significant amount on

average um when our when our data is in

about the hundreds of thousands here.

So we are um we have a significant

amount of error 50,000 relative to the h

to our units that our our data is in.

Right? Um so this score is not very

good. Um

and so we see that from the cross

validation. So look how easy the cross

val is. Again we just do cross val

score. We put in our model. We put in

our data. We put in our cross validation

uh strategy here which is k-fold and we

can generate these metrics across all

the fold combinations. So it's this

function is taking care of rotating

those and doing every combo with just

the 10 different combinations here of

the of the folds.

10 different instances where you have

you know 10 different folds are the ones

that are left out for evaluation.

Um so it's managing that for us using

this data right using this training data

here. Um and we uh we generate these um

generate these scores.

Okay. So that's kfold. It's not hard to

do. All you have to do is um just use a

cross file score. And we could change

this to mean squared error. That's you

know we could do that too. That'd be

pretty easy. Um, so that'd be no issue.

We just happen to be using the absolute

error here. Of course, we could use

squared error.

Were you guys able to get this to run

kfold scores?

It produces an array of 10 10 different

scores, which should make sense because

those are these are the um we're

splitting our data into 10 different

folds,

right?

10 different folds and leaving one out

to do our evaluation on. So the one that

gets left out every time is what's

producing these scores. So it's 10

different ones get left out when we

rotate through all the combinations.

And so we average these scores

and we get this amount of we get about

50,000 in error on average.

Um, what do you think would be what do

you think would be acceptable? So, if

our if we're predicting the price, like

if we're a real estate agent and we're

predicting these prices and they

typically are

Yeah, close to zero would be great.

That'd be fantastic. Closer to zero

would be better. The average is um

206,000.

So 50,000 is a decent percentage of

that. Um so you know you can compute it

as a percentage right? So 50,000 is a

decent percentage of that. Um probably

you want this to be less than 20,000

would be about 10% error. 20,000

right? So maybe like 30,000 somewhere in

there.

Yeah. 10% would be 5% error. 10,000

would be 5% error. That's true. That's

true. So that would be that would be

much better. So being closer to zero,

like the smaller the better, of course.

Of course. Um but yeah, I would say an

acceptable percentage of error is

probably 20%.

Probably 20%, which would be um like

40,000 or less would probably be

acceptable.

Usually when we usually when you build

models um 80% accuracy is usually uh

considered decent.

Usually considered decent

80%. So I'd say 40,000 or less would be

kind of ideal.

Does that make sense to answer the

question?

That's a good question. What value is

acceptable? I think probably less than

40,000 would be ideal. That's right

around 20% error.

All right, so that's kfold. Um let's do

just a regular hold out now. So this is

just using our training and test data.

Um doing model.fit and calculating an

MSE on the test data. So this is this is

just the um hold out strategy here where

we just have um this is less robust but

it's a lot quicker to do and easier to

set up. Right? So um this is using the

hold out strategy. So just a regular

um train test split.

Are we going to rebuild the model? No,

not necessarily. There's some things we

could do most likely. And like one thing

we did not do was scale our features.

Remember I said that's a pretty

important thing to do is to scale our

features. We did not do that. So that

would be an enhancement to this that

we're going to So I I actually do think

we'll do that later. Yes. So I think we

will actually do that now that I'm

thinking about it. Yes. One of the

things we can do is scale these features

using like a minmax scaler, a standard

scaler. that's actually going to help us

um that's going to help us do better

predictions.

So that that's one thing we could do. Um

but yeah, we will we'll try to see if we

can get better.

It should help it. Yeah, usually you

want to scale you want to scale the

data. That's something we didn't do in

our preparation step. We did a lot of

the things we should do. We removed nles

and we did one hot encoding to the

proximity feature like this one. Um

those are good to do but we didn't scale

any of these other we didn't scale any

of the features right we didn't scale

any of them. Um it you it will have an

effect. It usually when we scale it

it'll be a better model.

It'll it'll learn a little bit better if

we can scale the data. Um so that way

like these

um like ages aren't you know drastically

different than like in scale then total

bedrooms or income

uh those kind of things. So we usually

want these to be in a similar scale

range.

So we'll we will I think we'll scale

them coming up in a bit and it should

help the model.

We've talked about that before, right?

scaling usually is a good idea to do

when you're prepping your data for

modeling.

No, you want to you want to scale your

test data as well. You're going to do

both. You're going to scale your

training data, you're going to scale it.

So, that's actually a good point you

bring up is any transformations you do

on your training to build your model,

you should also do on your test set so

you get an applesto apples comparison.

You should always do the same

transformations.

Yes.

Would scaling data impact K? Yeah, it

could it could make it better. It could

uh yeah, it should impact it. We should

get a better model. So when we do the

different folds, we'll get different

we'll get better scores. Yeah, it it

will impact

uh yeah, if they're so that's a good

point. If they're going to use our

model, then yes, they have to scale the

data as well. If they're gonna if we

build the model on the assumption that

the input is scaled, then yes, they have

to also scale their data when they're

using it with our model. That's true.

I mean, not really. I'll show you why.

There's something that's actually going

to make it easier um that that will

automate doing the scaling for them. So,

they don't they don't have to do the

scaling manually. it'll just it'll

happen automatically when they use the

model. I'm going to show you something

that's going to automate that which is

going to be called a pipeline.

So that part will be automated and they

won't have to do that. So it won't be

heavy on the user. No, in theory it is,

but

has a really helpful tool to make it

easy to do that. So I'm going to I'm

going to show us that um later on in the

notebook.

No, the data data is not for a single

house. It's for like a neighborhood. So

there's a certain number of households

in the neighborhood and this is the

we're predicting the median house value

of that neighborhood.

Yeah. So there's a there's certain

number of households. There's there's

like an a median income, a population,

certain number of people that live

there. Um proximity generally of where

that location is. It also has a latitude

and longitude.

So,

and a median age in that neighborhood.

So, yeah, it's not just a single house.

Okay, let's go back to this was the hold

out strategy. So, this is a lot simpler.

This is just model.fit, right? This is

just model.fit on the training uh data.

And then we um can predict on the test

features and generate test predictions.

And then we can compute our error on

those um we can compute our error

amongst the test predictions and our

test uh label. So that's our useful mean

squared error function, right? To to

compute the MSE. Um let's see what the

MSE is. So MSE is right here.

Um now what we can do is we can take the

MSE

and we can take the square root of it.

So let's actually do that. Let's um do

MP. square root of the

um test

MSE

and we get um 67 we get 67,000.

So that's pretty high on this. So when

we just now look at the difference of

that, right? When we just do a train

test split, um

when we just do a train test split, we

get a worse score because it's not as

it's not as robust, right? We're not

showing that to many of the other uh

folds. So we get a lot more error this

way on the test data.

So this is um actually worse performance

just doing the train test split.

This is a really higher.

Yeah, we can. We can. I'm going to I'm

going to show us how to how the scaling

will be done automatically. Yes, we can.

Um there's there's a really easy tool to

do that will scale it automatically.

It's going to be later in this notebook.

I'll show us it.

All right. So, just to recap this, this

is fitting the model.

This is fitting the model. This is

making the predictions, right?

Model.predict.

So, this is making the predictions. And

then this is calculating the error, the

mean squared error, which is looking at

our test labels versus our test

predictions, right? And this is

computing the distance, the average

distance away from these values to these

values,

right?

And then we can also compute the R squar

R R 2 and we see that it's not a very

good R squared. 65 uh is not a very

great model

um because it closer to one would be

better. So this is still this is not

very good.

We know that we knew that from the cross

file score. But this is just doing um

this is just doing a hold out uh where

we do a train and test split, right? So

it's a little bit simpler, but it's not

quite as robust. Um

it's not quite as robust as the cross

valve, but it works. Um it's, you know,

we can do hold out. Um,

we can do hold out uh to to quickly

evaluate a model and see if we need to

make any adjustments.

It's a little bit quicker to run.

Okay. And any questions on it? Does it

make sense what we're doing here?

Model.fit to train it predict to get our

predictions. Um, this is pretty

standard, right? To train is the

model.fit it and then to use the model

to predict. We predict on the test

features.

Um, so this is passing on on all of our

features into this model to generate

predictions for every row. That's

something I also want to point out that

may be a little bit confusing is this is

a data frame. So we're passing in a

bunch of rows of features with columns,

right? So um, we're passing in a bunch

of data that looks like this. And what

we're doing is essentially making a

prediction for every row. So this will

generate a prediction. This row will

generate a prediction. This row will

generate a prediction. And on and on and

on. So this this predict will predict

for every row. And so we end up with

this collection of predictions here for

each row. And we're comparing those to

the labels that we have for those rows

from our from our supervised learning,

right? From our data set.

So that's truly supervised learning,

right? We have the examples and we're

comparing those to what our model is

predicting to to get our performance.

All right.

So let's uh let's try the other just so

you can see it. The leave one out. Now

the leave one out cross validation is

going to actually work the same way

where we put in the leave one out um

strategy inside of the cross file score.

Now here we don't need to specify how

many folds there are because we know how

many there going to be. It's going to be

the number of data points, right? So

which is actually going to be quite

large because there's 20,000 rows. So

this is going to be extremely

uh extremely um intensive because we are

doing um you know 20,000 examples and

leaving one example out to be our

validation and then um doing that across

every 20,000 uh examples.

So we could do it though just to see how

it works. Um we have this again leave

one out. We generate our cross file

score from our model our data and then

same scoring that we had before and but

this time we change our cross file to be

instead of our k-fold object we have our

leave one out object which is this

um and then we could run this. We can

compute our average uh across the all

the folds. Now this is going to be a lot

bigger of an array. It's going to be a

20,000 size array and we're going to

compute the average across it.

So, let's do that. It's going to take a

moment because there's lots. So, if you

notice it when you run, it's going to

take a little bit of time to run because

it's running across all 20,000 examples

and leaving one out. So, you have 20,000

and then one left out to uh test

against. So, it's quite intensive. You

can see it's taking a lot more time.

It's still running. It's taking a while.

Okay, just let that run. Still running.

So, if you guys try running this, it's

going to take a little bit of time.

Hopefully that makes sense why it's

taking so long, right? It's because it's

instead of doing 10 folds, it's it's

putting every data point but one is the

training set and then iterating through

all 20,000 points.

This takes a while to do.

Let's see what our

RAM our memory is a little increased.

Okay,

still running. That's okay. I'll let it

run.

Come back when it's finished.

Yeah, exactly. This is a This is for

This is giving us a performance

evaluation. This is like the average

error across all of our uh different

folds. Um now this is the extreme case

where we have the number of folds equals

the number of points.

Right? So it's an extreme case but yes

it's just like kfold. It's giving us

that performance estimate.

Okay. It's about the same right. This is

still around 50,000.

Not much difference, right? Still right

around there. But look how much longer

it took. That took 2 minutes to run. The

other one was pretty instant, right? So

this this took about 2 minutes to run.

So um definitely uh

yeah, definitely don't want to run this

uh too often. I think that it's

generally preferred to do k-fold if

you're going to do cross validation.

Generally want to do k-fold or just the

regular hold out train test split. Uh

generally better than doing leave one

out. It's just going to take too long

and um it results in about the same kind

of score as the kfold.

Okay,

any questions about um the cross

validation that we just did.

Okay,

good. And as it says here that the

stratified kfold is usually used for

classification. Again, we're not doing

classification yet. That's in going to

be in lesson four. So, we don't need to

worry too much about that. Just for

regression, um regular kfold is

preferred, right? Because we don't need

to um worry about distributing

categories amongst our folds uh in any

regression problems.

And as we see the error is kind of high.

Um there's going to be some things we

can do to improve that which will be uh

later on we'll learn about some more

advanced models. This signals that the

performance is bad. We probably need a

more complex model. Um one thing we

could try before we try a complex model

is to do scaling. We will try to do

scaling. I'm going to show us how we can

do that coming up um in a in a nice

streamlined fashion. Um, but uh outside

of that, if we still had bad

performance, we would likely need to use

a more advanced model. And we'll learn

about more advanced models uh in the

next lesson. And what's great is some of

those advanced models can actually be

used for regression. So they have

variations that can be used for both

classification and regression, which is

pretty cool. So I'll point those out

when we get to them. Um, okay.

So what I want to talk about now is a

way we can combat overfitting. So if we

have overfitting which remember that is

the case where the uh the we see good

performance on the training data but

then um it doesn't generalize over to

the test data. We get poor performance

on the test data. Um there's there's a

drop off there. Um that would signal

overfitting.

overfitting

and one way of um combating overfitting

is to do something called regularization

which we're going to talk about next. So

the key idea in regularization

is to

change our uh the change the way we

train. Essentially, what we're going to

do is modify our training

uh error function or sometimes called

the objective function or loss function.

We're going to change that to add a

penalty to penalize excessive complex

complexity. Essentially the the way that

we're going to penalize is by making

sure the size of the coefficients

doesn't grow too much which should

mitigate overfitting because remember in

linear regression what we are learning

are the coefficients right we're

learning the beta 0 the beta 1 the beta

2 and on and on however many betas there

are beta n we're learning all of those

guys um through the regression error

function we're trying to minimize that

error function. That's how it trains. We

talked about that on Monday.

Um so what we're going to do is um

basically penalize the these guys

growing too big and making sure we kind

of keep them small so that no one

coefficient has a dominant uh effect on

the model. And this should help with

overfitting and complexity. should make

the model simpler because all the

coefficients are going to be encouraged

to be smaller. They're not going to grow

too big. Um and this this has the effect

of making the model so basically make

the model simpler.

Make the model simpler is what these

regularization techniques are

essentially trying to achieve is is

remove complexity, make them a little

bit simpler, make these coefficients

smaller so that you can generalize a bit

better and and prevent overfitting. So

we want to prevent

uh overfitting,

right, is what we want to do. Um so

there's going to be a penalty and I'll

show you where that penalty gets added

and kind of what it looks like.

Um but uh to control the level of that

penalty we are actually going to

introduce another parameter to our model

um called alpha.

Alpha is going to scale the penalty. So

if alpha is really high that imposes a

stronger penalty on the coefficients um

which will make the model a lot simpler.

So the higher the alpha the simpler the

model we will get and we the the risk

with that is we actually underfit. So if

alpha is too big we may underfit the

training data

um a bit too much because it will make

the model way too simple. Um and I again

I'll show you what this means

mathematically in a moment. Um but on

the other hand if we have a lower alpha

this will have a lower penalty. it's a

weaker penalty term and that'll lead to

a model that is um a bit more complex.

Um which could um risk some level of

overfitting. Um so there's so there's

still the risk of overfitting if you

have a low alpha. And of course if alpha

goes all the way to zero, there's no

penalty at all. So you're back to your

original linear regression. Um which

could risk a lot of overfitting,

right? So you you generally want to pick

an alpha um effectively and actually

we're going to see h what's the best way

to pick alpha. Um we're actually going

to learn how to do that. I'm going to

show us how doing some tuning techniques

to pick what alpha should be. Um but um

a a pretty industry standard alpha that

most people default to is alpha equals

to one. So just just one which signals

that there should be some penalty. we

just have alpha equal to one is a

standard penalty. We don't want it to be

too high. We don't want it to be too

low. Like we don't want it to be a

fraction. Um but a penalty of one is

usually uh good enough.

Okay, I'm going to show you where that

comes into play in a moment.

Um but the whole purpose of doing this

is to mitigate overfitting, right? Um

that's what and and doing this penalty

is is called regularization. So adding

so going beyond just regular linear

regression adding this extra penalty to

to the training process um to penalize

large weights large coefficients

um is known as regularization.

Okay. Um and there's two common

penalties that are added. Um so there's

actually two different variations on the

penalty. Um we're going to study both of

them and um they're they're known as

lasso. So if you take linear regression

and add a particular type of penalty,

it's known as lasso. If you add another

type of penalty, it's known as ridge

regression. We're going to study both of

those and what their differences are.

But these are the primary two

uh regularization tech uh models that

are used um to take a regular both of

these take regular linear regression and

just modify the training process a

little bit in different ways. Two

different ways. um using that alpha

um to penalize the terms in slightly

different mathematical ways. So we're

going to learn about these two guys.

Lasso regression there. Both of these

are just offshoots of linear regression.

So underlying model is still linear

regression. It just adds different types

of penalties to the training process.

So both of these are still in the family

of linear regression. In fact, in um in

scikitlearn, they both come from they

both are still from the linear model

family in inside of the linear model

module, which is where linear regression

comes from. So there's still linear

regression. They just have different

styles of penalties added to them. Um

which we're going to see.

Okay, so just to recap that

regularization is the process of adding

a penalty to the training to discourage

complexity. In this case, we're going to

discourage large coefficients.

And um this should help prevent

overfitting.

And so uh these are going to lead us to

two different offshoots of linear

regression that have two different

penalties.

lasso and ridge regression, which we're

going to uh study next,

but they they function the same way as

linear regression. They will just have

different penalty terms added onto their

training process um to discourage

uh discourage um again those large

weights.

Okay, any questions about regularization

before we first look at our we're going

to look at our first uh variation on on

our first regularization technique which

is going to be called lasso regression.

Okay, let's look at lasso regression. So

what is lasso regression? It's actually

lasso is short for least absolute

shrinkage and selection operator

regression. Um and this will function by

adding a particular penalty to the

linear regression model. So again it's

based on linear regression. That's the

underlying model. It's just that during

the training process we are going to um

add a penalty which has the effect of

shrinkage of the weights. That's why

it's called shrinkage. It encourages

smaller weights through that penalty and

it also will shrink some of them so much

that they'll become zero and so it has

has an effect of kind of selection which

means that some of them get wiped out to

zero

and this means that whatever is left

over is kind of what's selected as our

features because the other ones will

have zero weight applied to them. So

this penalty will really favor small

weights um and penalize really large

weights. In fact, it will favor small

weights so much that some of them will

actually um be shrunk to zero um during

the training process. And the ones that

are left over are the ones that um are

the ones that are what we call selected

because they are the ones that remain in

in the training um after the other ones

get uh coefficients of zero. Um now when

you make some of the coefficients zero,

you are inherently making the model

simpler, right? There's less features

involved in the prediction that or less

features that have an effect on the

prediction. So this definitely makes the

model simpler. This lasso, this

shrinkage and selection uh process makes

makes the model simpler for sure. Um

and this is supposed to reduce

overfitting, right? If you make the

model simpler, it's not as complex. It

has less of a chance of memorizing

training data and not generalizing over

to test data. So our whole goal with uh

regularization is to make our model

better at generalization right over to

test data from the original training

data.

Um so how does this happen? We have to

go back to the

uh training process. If you guys

remember I I wrote out this equation a

little bit earlier which is the

distance. This is the sum of squared

distance between our labels and our

prediction.

This is basically the mean squared error

uh calculation that we're trying to

reduce when we build our model using the

training data. Um so this is just in

standard linear regression. This is the

um uh sum of squares uh distance right

so this is this is what the model is

trying to minimize when it learns these

coefficients.

So when it learns these coefficients,

it's trying to minimize this guy.

Minimize. It's trying to find the betas

that minimize this quantity.

Mathematically, that's what it's doing.

Um, and there's there's a algorithm that

will discover what the best betas are

that actually minimize this. That gives

us the line of best fit, right? That's

what we've been talking about for

regression.

Now in regularization

here's by the way here is that same

thing but we've just inserted our model

for the predictions. This is our model

just a fancy way of writing down our

model right it's the beta 0 plus all of

these betas. So beta 1 x1 plus beta 2 x2

plus on and on and on. Right? That's

that's what this uh means if you're

unfamiliar with the sigma notation. It

just means sum. So it's the sum of all

these guys or this term. Um so this is

this here is just a regular linear

regression

uh training regular linear regression

training. So we the training process

solves for these parameters right it

solves for these weights. We discover

what those are by minimizing this

quantity. That's the whole training

process. Um but when we do lasso

we add a penalty which is this

here is our penalty.

So basically um we take our linear

regression training which is this and we

add on a penalty which is this and you

can see exactly what this penalty when

when you minimize this penalty it's when

these weights are small. So this

encourages

So minimizing this quantity encourages

small weights

encourages small betas

beta I

right you or in this case beta j sorry

this encourages small beta js uh because

we want this thing to be minimized

minimize

So um what's going to make this minimal

is of course the line of best fit and

small weights right are going to make

are going to bring this error down the

most.

So um and here's our alpha right here's

our alpha. So you can encourage a higher

penalty with a larger alpha or a lower

penalty. If alpha equals zero

what happens to that term? It just goes

away. So if alpha equals zero, there's

no penalty and we're back to uh we're

back to regular linear regression.

We just have regular linear regression

because we have no penalty at that point

when alpha equals zero. So the smaller

alpha is, the less penalty we're

enforcing in in the regularization.

Okay.

Now what happens is in reality when you

train with lasso. So this is lasso is

this particular penalty. This is called

the lasso penalty

or sometimes um people call this the L1

penalty.

Um L1 just comes from the fact that this

is the first power or absolute value. Um

so it's not a squared penalty. It's a

single uh single power penalty

um there. But when you add this lasso

penalty, what can happen is it c it does

because the because you're minimizing

this, it does encourage some of these

weights to become zero.

So some if you're really trying to get

the lowest quantity of this,

the lower the better.

What makes this thing lower is of course

if some of these go away. If some of

these go to zero then that of course

will lower this as much as we as much as

possible. Right? So what happens during

the training is some of these

coefficients actually they're encouraged

to be small because of this penalty. But

some of them will actually becomes will

actually become zero um in order to get

the best model the best fit. Some of

these will actually get so small that

they'll basically become zero. And that

means that that that feature basically

has no effect anymore. It's it's been

the model has been simplified, right?

That feature no longer really has an

effect.

So just to call out the alpha again um

if alpha zero some co uh basically you

have your linear regression you're back

to linear regression because alpha 0 is

just wiping this out and you're back to

linear regression. Um if alpha is

infinity now if alpha is infinity that's

an extreme. So if alpha is infinity the

only way to make this minimize is if all

your coefficients are zero. If every

beta is zero, then this will lower the

the error as as much as possible. So you

basically have no model. So if all

coefficients are zero, you have no model

and that's useless. So you don't want

your penalty, you don't want your alpha

to be huge is what this is saying. You

also don't want your alpha to be small.

You're basically back to linear

regression. So you want something in

between. Um and the typical typical

value is alpha equals 1.

typical is alpha equals 1

to have some level of penalty there. So

just a regular kind of regular penalty

term.

But we are actually going to have a way

to test and evaluate which alphas are

the best.

Um,

basically you can, yeah, you can have a

you can have a penalty that's close to

zero. You can get rid of this if just a

regular linear regression performs

pretty well. You can basically have no

penalty in that case.

Yeah. So near zero or like it could be

that adding a little bit of penalty

actually helps the overfitting and it

could be really small. One thing that

we're basically going to do is have a

strategy to try out different alphas,

try different alphas

and evaluate performance

and then we can decide which. So that's

what we're going to do is have a

strategy to just plug in different

alphas, generate the like train the

model, and then see what its performance

is and see if those alphas are good.

What what which alpha is the best? We

can evaluate that

because we can train the model and see

what it performance is,

right?

Yeah. Yeah. So we'll do that. We'll

practice that.

Okay, great. Any other questions about

this lasso regression? So, remember this

is linear regression here. This is the

this is how you're training to find the

betas in linear regression. So, this is

just linear regression uh um training

function there.

We're adding a penalty which is this is

the lasso penalty

lasso penalty there right we're adding

that this is known as regularization

and the goal of regularization is to

prevent overfitting so you add a penalty

here this makes the model simpler which

prevents overfitting it helps you

generalize better when it's simpler Any

questions conceptually on this? We're

going to do a code example with it

coming up, but any questions on this?

Uh yeah, you you so that's the thing,

Ronald, is you may be willing to

sacrifice some accuracy in order to

generalize to unseen data because

remember that's what we're really trying

to get after is we may be willing to

sacrifice some accuracy on this training

data in order to have it perform better

on the test data, right? We may be

willing to do that. That's a willing

that's an okay sacrifice

as long like if if it generalizes

better. That's what we want. That's what

we're trying to do here is add a

penalty, make the model simpler and help

it generalize better to new and unseen

data. Right?

That's that picture I've been using with

the with the um train and test split.

Where is the square?

So in the model there's no square. So

remember the model is the model is this

um equation uh that has no squares in

it. Right? It's beta 0 plus beta 1 x1

plus beta 2 x2 plus beta n xn.

That's the that's the linear regression

model. This is the now this this is the

model but this is the equation that

helps us train and find the betas. This

is how this is what we find the betas

with. So we'll continue. Um we were

talking about the lasso regression which

uh adds it takes linear regression right

which is this optimization and adds in a

penalty um scaled by the alpha. Um, and

what that does in order to minimize this

whole thing, it encourages these to be

small uh as possible. Um, which makes

the model simpler, right? The weights

don't get overly big and complex. Um,

they they tend to stay small. In fact,

some of them can even go all the way to

zero. Um, which makes the model even

more simpler,

right? Um, so let's practice uh using it

in code. It's actually really easy to

use. It's going to be essentially the

same uh style and and code as linear

regression except we are um just going

to have to uh put in our alpha parameter

um when we use the lasso. So here we are

um from the linear model family right

which makes sense. It's a linear

regression offshoot that has this

penalty in it during the training. um we

are grabbing our lasso regression. Um it

also has a version of the lasso that

we're going to take a look at that is

used for cross validation which is

really um convenient as well. So it has

a cross validation lasso which is a

really convenient um combination of

basically cross val score and lasso um

all in one. So it actually is really

nice to use that way. Um so we'll take a

look at that example. Um, but we are

importing it. The main thing is going to

be the lasso model here. Um, we're going

to be using a different data set for

this one. So, not the ocean uh data, but

this hitters data, which is a baseball

data set. Um, so it has 322 rows um with

20 different columns and it looks like

this. So, you want to download that one.

Um, hopefully you guys have access to

that one.

Um,

so I will upload it into

this.

So give me a moment.

There's that. And then we can run this.

Okay. So we are displaying the data and

so it has um the the hitters names and

then it has a bunch of different

statistics. These are all baseball

statistics.

Um, if you're unfamiliar with with them,

that's okay. It's not a big deal. Um,

but just different baseball stats here.

Okay. Were you guys able to load that?

Um, if you're following along, were you

able to load that? You should have

access to this data. The hitters CSV.

This is the one we're going to use for

the lasso model

to build a lasso model.

Yeah.

Okay. Able to load that one. Perfect.

Okay. So, able to load that one. Um, and

we take a look at the the head. Um, so

we're actually going to uh drop this

unnamed column because we don't care

about their name. it's actually just the

batter's name, which is not going to be

useful in modeling. Um, so and remember

that's generally true like an ID, a user

ID, like a customer ID, a name, that's

usually not going to be useful in any

kind of modeling. So we're actually just

going to drop that uh column and we're

going to do it in place.

And access equals 1 means we're dropping

that column. Um, so we're going to drop

that and we should no longer have that

column. And we have all of these guys

now. So you want to run that. This will

drop that. Um this will drop drops the

column in place.

Um and now we can see we have uh all we

have this data where um we have this

data where it's now removed. So this

that column is now gone and now we have

these guys. Um, do you notice anything

about this

from the info?

Looks like we have a couple categorical

features, a few of them, league and

division

and new league. What do you notice about

this?

Nolles. Yep. So, there's definitely some

missing data there um that we're going

to have to deal with.

So, it looks like there are uh there are

59.

Um there are 59. Now, we could we the

alternative to doing that is we could uh

we could just use our usual code where

we do df.is is uh is null.

Um and then we do uh dot sum to total

those up across our different columns.

And we can see that uh we have 59 of

those in this salary column. That's this

is the standard way of doing that,

right?

Standard way of doing that. And we have

so we have 59 of those

59 of those. So, we have to deal with

it. Any ideas on how to deal with it?

Any ideas on how to deal with it? This

is now This is 59 out of 300.

So,

what do you guys think about that? It's

a little bit different than 200 out of

20,000. A little bit different. We have

We have about 60 out of 300.

There's a decent amount.

Any ideas on how to handle this one?

Replace. Yep, we should replace. What do

you think we should replace with?

It's a float. It's a floating point uh

value.

By the way, something unique about this

that's a little different than usual,

too, is that the uh this is actually the

column we're going to use as our label.

So, we're actually going to predict the

salary based on the uh based on the um

rest of the features. So, we definitely

need to fill in these nles, right?

because they're actually going to be the

labels

and we're missing some labels uh in our

data. We we definitely need to fill them

in. Yeah. So, we're going to replace

them.

All right. So, we'll we will replace

them down below. That's going to be

coming up. Uh we'll come back and

replace them. um before we replace them,

we're actually going to get our uh one

hot encodings for those three different

um features we have. Um so we do uh get

dummies with this. Now um of course we

don't need to do this if we just so this

code we don't need to do if we just pass

in the dtype here

um which is uh then we don't need to do

this. So we can comment this out.

Um so now what I want you guys to notice

is this is the alternative to what we

did before where we are purposely just

doing these columns not the whole data

frame but just doing these columns and

then we can um concatenate those these

one hot encodings. We're going to

concatenate back to the data frame.

Right? So if we do our dummies and then

do dummies.info info. Um, we can see

that we end up with six new columns. And

in fact, we can do dummies.head

and take a look at what those are.

Right? So, these are league A, league,

uh, N, division E, W, division W, new

league A, new league N.

Okay.

So, um these are uh these are our one

hot encodings for these three different

features which are strings, right? So,

those those features were strings. If

you go back up, those were our only

string features we had. So, we've one

hot encoded those so we can use them in

our model. What we need to do is just

concatenate this back to our data frame.

Right? So, we just need to concatenate

it back into our data.

Okay. So, what we're going to do then is

we're going to grab um we're going to

grab Y as our salary. And of course,

we're going to fill nles on that Y

coming up shortly. But we're going to

grab Y as our salary and X new. Now

before building a full X, we're going to

take a look at X numerical as our data

frame minus these columns. The reason

we're doing minus those is because we

are going to concatenate our dummy

variables back into this that are going

to replace these guys. So we're going to

replace these anyways with our one hot

encodings. We don't want the strings. So

we're going to get rid of those. And

we're also going to get rid of the

salary because that's going to be part

of our that's just the label. So we

don't want that in the X, the eventual

X.

Are you guys able to run this one?

Hope I'm not going too fast. You guys

able to run this? And does it make

sense? What we're doing is we're putting

our labels in Y, which is what we

usually do. So we're going to predict

the salary

and we're getting ready to build the X.

But before we first want to get rid of

those one hot the the strings. This is

getting rid of the strings

and this is getting rid of the label and

that's going to be part of our features.

What we need to do is build our final X

by concatenating our dummies with this.

Do you guys see that? We're going to

concatenate our dummies with this to

build our final X.

But but prior to doing that, we need to

get rid of these string columns here. So

we're dropping those

dropping those from the uh data frame uh

and getting a numerical uh x numerical

here.

You can see the columns of that are just

these guys here. So the the results we

need to concatenate our we need to

concatenate this guy um into this and

then that'll be our full x all of our

features.

Okay. So you can see x is going to be

pd.con

of this with our dummies.

This with our dummies. And um

uh instead of doing this, I'm actually

going to do the full dummies. We don't

need to um pick just a few columns.

We're actually going to do our full

dummies here and um do x equals 1. Now,

the reason that's the case is because um

this will get rid of one column per

feature and basically assume that if you

have a if you have a zero, the other one

should be a one. If you have a one, the

other one should be a zero. Um so it

basically makes that assumption because

we only have two of them. Um so whenever

there's a one, the other should be zero.

Um, so you can get away with just having

these three, but um I think it makes

more sense to just have to have the full

dummies,

but by process of elimination, you can

get away with just using two of them

because anytime you have a zero, the

other one should be the other feature

would have been would have been a one,

right? And vice versa, when there's a

one, the other feature would have been a

zero.

So we do that one.

And you can see all of our uh all of our

one hot encoding features end up back in

there.

So this is the code that I want you guys

to run. I think it makes more sense. It

follows along what we've been doing.

um which will concatenate our dummies

back to our features here to build out

our full X. So now X is all of our

features. Um remember X

X X contains all of our features

now.

So X contains all of our features and so

we have all of this now.

Okay. Were you guys able to run this

one?

D. We have y, we have x. We still need

to deal with the nles in y. So that

something we still need to deal with.

But hopefully you have this. Now

all of these are numerical.

So that should be good with the model.

That's one thing about X is you should

you our X should have all numerical

features, right? Because it's going to

go into a model to to learn those betas.

So it needs to have all numerical

features,

right? These are going to be all

numerical, which makes sense. We change

we did one hot encoding to change all

those guys to numerical.

Sorry, I'm scrolling down.

Okay, we do fill in the nator. Okay.

Okay.

Any questions so far? So, we're just

getting our data ready. We haven't

applied the lasso yet, but we're just

doing some prep. Now, hopefully you guys

recognize th these are some standard

steps that we're taking when we do our

modeling. We have to do these data prep

steps. They're necessary. And so, if it

seems like it's a lot of work, that's

because it is. It is work that you do to

prepare your data to get ready for

modeling. You have to do that. Okay.

So, we're doing that here. Um, now we're

going to do our train test split because

we're just going to do uh we're going to

do hold out here. So, we're doing a

train test split with about with a test

size of about 0.25. So, again, anywhere

between 02 to.3 would be okay.

Um, so uh it's our choice. We could do

02. We could do 3. We could do anywhere

in between there. We're doing 0.25.

That's fine. Um, that's okay. So, we we

build our train test split right there.

Um, so pretty pretty simple and we've

seen that a bunch of times with our X

and our Y data frames. There we have our

train test split.

Okay.

Um, now what we're going to do is do our

our scaling. So, we're we didn't do this

last time, but we're going to do this

now as uh because we should get in the

habit of doing that. Um is um we're

going to um go ahead and scale our

features and we're going to use the

standard scaler here uh to do that

scaling. Okay. Now, we could use minmax

scaler that's fine, too. We're just

going to use the standard scaler here.

Um and remember we are going to uh um

use the standard scaler from sklearn and

we're going to transform our features uh

uh according to our um according to our

training data. So we have our

pre-processing standard scaler here. So

we import that guy and then we um build

our standard scaler and fit it on the

training data only on the numerical

features. Um so that's which is going to

be uh all of these guys. So we're doing

the scaling on all of these guys. Now,

something to note is that we are not

scaling all of these one hot encodings

mainly because it doesn't make sense to

scale those really. They're zero or one.

They don't need to be scaled, right?

They're already zero and one. So,

they're they don't need to be even if we

were doing minmax scaling, it's going to

put them between zero and one. It

wouldn't affect it really, right? So,

these one hot encoding features, we're

not going to scale because they're

they're always going to be zero or one.

There's no need to scale them really.

Um, but we're going to scale all the

other features here that are floats.

So that's these guys here. These

numerical features we're going to scale.

Okay.

Don't need to we don't really need to

scale the one hot encoding. Uh, it's

pretty much already scaled.

Oh, you should change that. Um, go back

and rerun go back and rerun this. But

make sure you have your data type as int

here.

Make sure you add that in there to

change that over to integer and rerun

that and then rerun the rerun the

concatenation.

So make sure you run this

and then u make sure you rerun this and

rerun the concatenation part which is uh

this

Okay. So, we go ahead and fit the um

scaler to this data and then we're going

to transform our training features,

those numerical features. Um we're and

then we're going to uh transform these

features uh uh the test features in the

same way. So we're going to perform the

same transformation from the scaler on

the test data. So that's something

really important I want to note here is

that we always scale both the training

and test data. We always scale both. Of

course, we're going to train the model

on the training data. Um, but we are

going to also test it on the testing

data and it also needs to be scaled

because our model that we build is going

to assume scaled features. The

coefficients that it learns are going to

be assuming scaled features.

So we need to also scale our test data

in the same way. So we're doing that as

well.

So, we scale that. And now we have our

uh training and testing features have

been scaled.

No, we haven't replaced. We're going to

do that. We have not yet. We're going to

do that coming up in a minute. Yeah, we

haven't done that. Um, it is it is the

label. We definitely need to replace

NLES. We just haven't done it yet

because it's not in the features and

we're doing all of our uh uh

pre-processing to our pre-processing to

our features.

Yeah. So, we're definitely we need to

we're going to in a minute.

Okay. So, if you look at the data now,

it's all been scaled. So, these are all

um zcores. These are all on a much

better scale now. Um, and these are we

still have our one hot encoding features

which are zero or one. So this scaling

should lead to a better model than if we

didn't scale. So scaling is really

important. We can see that here.

Okay.

Now, um, let me ask you guys, were you

able to run the scaling? Are you caught

up to here? If you're following along,

were you able to run the scaling?

Okay, great. Great.

Awesome.

Okay. So, uh what we're going to do now

is replace nulls in the uh replace nles

by calculating the median of the data.

So, what I want you to notice is that we

are taking the NLES now this is um this

is on purpose is we are purposely taking

the NLES um out of the median

calculation. So we're skipping the NLES

when we compute the median because we

don't want those NLES to affect the

median calculation.

Um so we compute a median salary here

and then we fill our NLES with the

median salary um from the training data.

So this is our choice. This is a choice

um to use the median and it's also a

choice to use the training set median

for both train and test. What we could

have done, this is an alternative that

we could have done is use the entire

column and then um use the median of all

of the data to replace. That's really up

to us. Um this is one way of doing it.

We could have done before we did the

split. We could have um filled in with

the median earlier. We chose to do it

here mainly because it doesn't affect

the features. So, we could have done

this earlier and did it before we did

the split and filled the NAS. Um really

doesn't it's doesn't matter that much

which way we do it. Um but we do need to

fill in NLES. We cannot have those be

null when we when we put it into our

model. So some way we need to fill in

NLES. Um and so in this strategy we're

filling in our Y train um with the

median salary from our training data.

And same with this we're filling in with

the median salary of the training data

as well. But that's a choice f we could

fill in with the mean with the average.

Um we could fill in with the we could do

it with all the data together before we

split it. we could have filled in with

all of the the median across the whole

data set. Um either one works. You can

do it either way, but we we did it um

later here to show that it doesn't

really affect the features. So, we can

choose when we do it, right? It doesn't

affect the features at all. So, we can

do all of our pre-processing on the

features and then do our label uh

filling and nulls um if if we have them.

uh x numerical. Um make sure you're

running uh this

uh x numerical was defined here

when we split it apart um from

uh when we dropped these columns here.

So make sure you're running this. This

is x numerical

gets defined there.

So, go back up to uh this cell

where we split apart the y and we and we

have the x here x numerical.

Make sure you run this.

Make sure you run this. And then you can

run these. Then you run this to build x.

All right.

Are we up to here with this filling in

the labels?

Uh because then we can build our model

once we're up to here. We've scaled

everything. We filled in our NLES.

We've gotten one hot encoding.

Yeah, it is. That's why you know that's

why we spend a lot of uh time on model

on data preparation with pandas, right?

That's why we did all that pandas work

for sure. Yes, there is a lot of work

before we can build a model.

Yes, the mo do you guys notice that like

the modeling is relatively easy. It's

just a fit and predict. The modeling is

actually really easy. It's all the other

work that's that's more involved, right?

more code.

The modeling itself is really easy.

It's just it's just one line of like

ffit.

Yeah, pretty easy to do.

And then you do evaluation which is a

couple lines.

Yep. There's these are all the these are

the common steps. All these steps we're

doing are very very prototypical in

model building is you let's just go back

through this to see what we did right we

imported our data

um we analy we dropped this name column

because it's not useful to us so we

dropped that um we filled in the nles

eventually um but you know if there were

any nles in our features we would have

to deal with those as well by replacing

them or dropping the rows like we did

earlier Um

and then we do one hot encoding because

of course we can't have any string

columns going in our models. We got a

one hot encode.

Um we uh then build our X and Y by

concatenating the one hot encoded back

to the numerical features.

Then we train test split. Right? That's

pretty common. Or we could do cross

validation either way. Um the K full

cross validation. Then we scale. So, we

didn't do this last time, but this is

something we should get in the habit of

is scaling um our features. So, we do

that and now we're ready to model. So,

now we're ready to model. Um so, that's

this part.

Okay. So, let's do the model. Um the

model's actually uh pretty easy to do.

So, we're going to use a lasso. So, we

have a lasso model here. Notice what

we're setting our alpha to. So the big

parameter we really need ignore this

iterations. We actually don't really

need the we don't really need that

parameter. Um so just ignore it for the

moment. But the big one that we're

setting here is the alpha. So when we

did linear regression, we didn't need

any parameters to go inside the linear

regression object. We didn't need any

parameters, right? Because there are

really no parameters of it. But for

lasso, the important one is the alpha.

And so we need to know what to set alpha

to. Um let's start with alpha equals 1.

That's a good starting place. So a

typical um starting point

for alpha

um is uh is one. So that's a typical

starting point. And so we can set alpha

equals to one. This max iterations is

the the parameter that governs the

training process because it is

iterative. So if for some reason we we

can't converge to the right betas and

we've run it for 10,000 steps once we

pass 10,000 steps, uh it will stop and

just give us the betas at that point.

But it will likely never hit this

number. It'll converge before then. So

um we don't really need to um specify

it. So, I'm actually just going to get

rid of it. Um, it's not really a big

deal. It should converge before then.

Um, but if if we want to set like a

maximum step size in the optimization,

we definitely could there. Uh, but not

concerned about that too much. But

here's our lasso. And then we're just

going to do a fit on our data. So, look

how easy that is. Just like a linear

regression. Lasso.fit,

right? So, we do fit. Um,

oh, I didn't run this. I'm sorry. I got

to run this. Okay. Actually, that's a

good example of what happens when you

don't when you have nulls, right? So, it

says our our null contains nan. That's

because I didn't run this. But now, that

should be filled in. Now, we should be

able to run this. Okay, perfect. So it

runs.

Okay. So you can see what the intercept

is. Um this is one of our coefficients,

right? The intercept is 457. And look

now what's really interesting about the

coefficients is look at what some of the

coefficients are.

Some of them are actually zero, which is

really So some of them ended up being at

zero, which is very very interesting.

that means that those features get

cancelceled out and they're basically

not part of the model which is really

interesting. Um so we have all these

coefficients and some of them are zero.

Yeah, negative0 is just because of the

convergence like they started out

negative and worked their way up to

zero. it. Negative zero really just

means zero, but they just were coming

from they were like small negatives and

ended up at zero

during the training process. They were

negative at one point and ended up zero.

Um

so yeah, negative 0 just obviously means

zero. Um it's still still zero there.

So what's interesting is some of these

features ended up uh being zero which

you don't usually see in a linear

regression. So if we were to train this

using a linear regression we typically

wouldn't see that but some of these turn

out to be zero because again we're

encouraging those betas to be small.

we're encouraging them to be uh small

and so um you know what happens is some

of them can be shrunk all the way down

to zero meaning those features don't

contribute that's a really simple model

at that point right so we've taken

something complex that includes all of

these features and actually reduced it

into something simple that only includes

these features

right

so that's what it does um now we need to

evaluate this to see how good of a model

it is. But that's what this is saying

here in this text is that um a positive

uh coefficient indicates that as the

independent variable increases the

dependent variable also increases.

Negative coefficient means as the

independent variable increases dependent

decreases because it's reducing the

value. Um and lasso is known for feature

selection by shrinking some of them to

zero effectively removing those

variables from the model from the

equation right

um

so that's what happens

some of them end up being zero

were you guys able to run this this

lasso uh fit which is the training of

the lasso

No, it doesn't ensure there's no

overfit, but it helps with overfitting.

It's supposed to help by making the

model simpler. And this is definitely a

simpler model because it's removing some

of the features from the model

essentially, right? Because some of the

features aren't going to contribute.

It's a simpler model.

It doesn't it doesn't mean there's not

going to be any overfitting, but it

helps prevent it. That's what it's

designed to do, help prevent it.

Yeah. So, higher coefficient. Yes. The

higher coefficient means it's a more

important feature towards the

prediction.

Yes. That's what it means for sure. The

higher the magnitude, the more of a

contributor towards that prediction. Uh

it is. Yes.

And it's not just it's it could be

higher positive or negative there. Like

a higher negative is also a pretty big

factor,

right? So So you want to think about it

in terms of absolute value.

does not guarantee but helps. Yes, it

doesn't guarantee it but it's designed

to help overfitting, help prevent it.

Yes, absolutely.

Okay.

So let's do some evaluation. Um so let's

do in this case we are going to do our

predict

Oh, yeah. I'm not sure why that's the

case.

Interesting.

We could try increasing the um max

iterations.

Okay, that's why. Yeah. So then you get

that result with the with the higher max

iterations.

It doesn't get cut off there.

I think that's why you probably left

this in there,

which is fine. You get about the same

numbers.

Yeah.

All right. Let's evaluate this. So,

we're going to to to do evaluation. I

want you guys to see again, we should

get in the habit of doing evaluation,

which is taking our model and predicting

on the training and predicting on the

test sets, right? So we predict on the

train set and calculate our MSE

and we um calculate our R2 score um or R

squar score I should say. Uh but again

the MSE is the one we're really going to

use mostly. Um but we calculate so we do

our predictions and then we compare that

into our mean squared error with our

labels

and we uh go ahead and do the same thing

with the test. Right? So we do uh

lasso.predict

on our test features and we go ahead and

compare that with the test labels. And

so what we're doing there is generating

our MSE.

So we we take a look at our MSE and we

get uh 84,000

MSE. Um and so of course we could take

the um what we could do with that is

take a look at the um MSE on the uh we

could do um MP. Square root

and do the square root of the MSE test.

and we get um 340. So this would be in

the units of our label. So, if we go

back and look at our label um for some

of those um

so uh we are in 300s and our data is

like right around the 500. So, of

course, if we describe this um we could

see what the statistics are of it. So,

we could do df.escribe describe and

generate that. But that doesn't look

like a very good error, right? If these

are in the 400s, um that's that's not a

very good error.

So again, it's not a very great model.

But one thing I want you to see is that

it's it's not overfitting.

Um if anything, it's actually

underfitting, which is what this kind of

um MSE suggests, right? because our

error here is 84 uh excuse me 84,000.

Um

our our area here is 84,000,

excuse me. And on the test set it's

116,000.

Um so these two errors are both bad. So

it's not overfitting. This is actually

underfitting. So it's not overfitting,

it's actually underfitting. Um, and so

that's the risk with something like

lasso is that it's making the model a

bit too simple and we actually risk

underfitting, which is what happens. We

have too much error across both the

training and the test set. Overfitting

is when we do we have really good

performance on the training set, but bad

performance on the test set. We're not

overfitting.

um we are uh underfitting because our

performance is not good either way. Even

this R squar is pretty low. It's not

even at 50%.

Okay, so that's so we we do the

evaluation and again the evaluation just

comes down to making predictions and

computing our error amongst those

predictions to our labels. That's always

what the uh evaluation is going to be

for MSE.

What's the ideal MSE? What do you think

it should be? What is So, think about it

like this. The MSE represents the

average distance between our predictions

and the labels.

So, if we're getting it right all the

time, what's that distance going to be

if we're always right? What's our

distance from what's our distance from

our predictions to our labels going to

be if we're always getting it right?

Zero. Yeah, there's not going to be any

distance. It's going to be right. It's

going to be perfectly aligned, right?

There's going to be no distance there.

So, yeah, an ideal MSE is zero.

That's an ideal MSE.

So, anything close to like the smaller

the better for MSE. The smaller the

better. Um, for this R squared, uh, it's

it's a scale between 0 to one where one

is the best. So, one would be perfectly

aligned predictions. Um, so, and again,

this this is we actually multiply by 100

to get uh because it's it's a number

between 0 and one. So we get about 47%

which is not good.

Okay.

All right. Any questions on this

evaluation?

All right. I want to show you something

which is

Yeah, this that's true. the scale of it

matters on the data because we should be

you should always interpret your MSE in

the scale of

um your your labels because your labels

like in this case our labels um you know

we could take uh for example we could

easily let's actually do that let's take

the average

let's take the average of our labels on

the training data

and and we can see what those are. Um,

so the average is 500,

right? The average is 500. And look at

what our uh square root of our MSE is,

which is in the same units as our

original. Um, so we have uh quite a bit

of error. 340 when our units are right

around 500.

So that's quite a bit of error.

Yeah, MSSE of zero means our our uh our

predictions are nearly identical to the

test labels. Yes, that's what MSSE of

zero means. There's zero distance.

So closer to zero, the better.

But we talked about it as you you really

so the rule of thumb should be what is

your RMSSE as a percentage of your

typical value. So your typical value is

in the 500s. Our our RMSSE is 340.

That's just really high. That's over

like 60% of that value.

So that's just a lot. That's too much

error. What we would love this RMSSE to

be is under 20% of the typical value. So

that means on average we are 20% or less

off in our prediction. That would be

good. That would be pretty good. That

means we're like 80% accurate,

right? That'd be pretty ideal. So you

got to think about it in terms of this

RMSSE which is in the same units as your

labels.

This is the

RMSSE

which is in the same units as the

labels.

So and then to interpret this we have

340

is compared to

typical

um salary unit of 500

right so this is uh quite a bit when the

typical value is 500 and we are off on

average by 340 units

that's so much relative to the typical

value

that's just too. That's a lot of error.

That's not a very good model, right?

It's underfitting. It's definitely

underfitting.

Yeah. So, that's a great question. What

should we do from here? So, um because

we're underfitting

um we should use a more complex model.

So uh we're going to learn about those

in lesson four, but we should use

something different. This linear

regression is still too basic. Even with

lasso, it's still too basic.

Yeah, we're underfitting because we But

it could also be we're underfitting with

a regular linear regression. We should

test that out. Um, and maybe it would be

an exercise for you guys um to test that

out yourself. It shouldn't be hard to

do. Um, you already have all the data

scaled. You So, do you see how you would

do that? You would just come in here and

build a linear regression rather than a

lasso and dofit. And then you would

evaluate it the same way with a predict.

It's really easy to do that. And then we

can compare that um to to this. It

shouldn't be that hard to do that,

right?

And something you guys could do for

sure. Um,

is build the linear regression and

actually compare it and see what kind of

difference it makes. I mean, we honestly

we could do it ourselves. We could do it

right now. Maybe it's worth trying that.

So, let's build a linear regression

for comparison.

So we have our linear regression

uh is linear regression and then we do

ffit linear regression.fit fit

right so so this will train it um and

then we can evaluate it so lin mse is

mean squared error

and then we can do our um let's do our

training let's do the training and then

um let's predict

actually let me do that here

uh y prediction

train

linear

equals um linear regression.predict

and then we're going to predict on our

training features.

Okay, do you guys see what I'm doing?

I'm building a linear regression for

comparison.

I'm doing fit here to train it and then

I'm making some predictions on the

training set and we're going to evaluate

those. I'm going to replace that here

with y prred

uh train

linear. So these predictions

Okay. So, if you guys want this code, I

can paste it in.

So, let's see what the RMSSE for just a

linear model is.

It's a little bit better. It's better

for sure.

So 289 is better than this 340. It's

better. It's getting closer to zero.

It's still underfitting though,

right? And that's just on the training

set. Let's look at the Let's do the same

thing, but on

Let's change this. Let's swap this out

for um test

And then let's do test.

And then let's do test

test.

And then

test test.

Okay. Okay, so this is producing test

predictions on the test set.

We are generating an MSE test

and then we're doing MSE test

which is using the test labels and our

test predictions and then we take the

square root of that for RMSSE and then

we're going to generate that. So it's

still under fit. I mean this is still

high. This is still high um on the test

set and versus on the training set. So,

it's still pretty high. Um, even the

basic linear regression is under is

still underfitting. Still underfitting,

right? Even without the lasso,

which is lasso is supposed to help with

overfitting. It's definitely not

overfitting. Um, it's definitely

underfitting, but this is a signal that

it's kind of overfitting because this is

performing better on the training data

and then it gets worse on the test data.

Definitely gets worse, right?

Did you guys follow?

I'm just running this above I'm running

this above this. It doesn't matter where

you put it. We could uh we could move it

down.

We could move it down to I just ran I

just picked a new cell right here and

ran it. But we could move it actually

let's do that. Let's move it down

to

after the lasso evaluation.

Okay. So I just moved it there.

And then let's move

this down.

So I just put it here after the um after

this. So this is the um this is

basically the objective function right

of the training process. So during the

algorithm that runs when we call ffit in

scikitlearn it's going to find these

betas right it's actually going to learn

what these best betas are for our model.

Um this is our model here, right? It's

the combination of betas times our

features um plus an intercept beta. Uh

so that's our model. But um we penalize

those large uh weights in absolute value

by um adding a penalty term like this um

where alpha is some level of penalty

that we want to provide. Usually alpha

equals 1 is okay. But um actually what

we're going to learn uh to finish out

this section is there's going to be a

systematic way we can test out different

alphas um that represent the level of

penalty we want to uh apply to lasso or

even ridge

uh regression. So that was the lasso and

um if you guys remember using it was

super easy. Uh we worked through this

problem with this um baseball data um

and we had uh

let's see scrolling down we um split out

our numerical data and we did uh we one

hot encoded our our categorical data

combined it back together. Hopefully

that um rings a bell there. Um and we

actually scaled our data which is pretty

standard to do is we do some type of

scaling to our features especially our

numerical features right want to scale

those in some way whether it's minmax

scale or standard scaler um want to do

that and so we did that for this example

and then we um ran the lasso regression

which is pretty easy to use. You just

use the lasso object and you pick an

alpha here. Um, again, we are going to

have a way to test out different alphas

that could be candidates and we can see

which one's the best. Um, so I'm going

to show us that today coming up shortly.

But that was that was the lasso. If you

guys remember, we did that. Um, this it

we compared that to a basic linear

regression which is just this pretty

straightforward just a fit and then

predict and then we can generate mean

squed error. Um, still not a very good

mean squared error on this data, it's

still fairly large. Um, so it's still

not, no matter which model we use, it's

still not very good, but at least we can

practice doing that comparison. That's

what we did last time. We did this on

Wednesday.

Um

and then

we saw that the effect of different

alphas we had a lasso um

we had a lasso uh cross validation

example here. So beyond just using a

regular lasso model that um scikitlearn

has a lasso cv which allows you to try

out different alphas uh with cross

validation and um figure out what the

best alpha is. Um, now we're actually

going to have a different strategy

that'll instead of just picking random

ones, we can actually um supply multiple

parameters that we may want to test. Um,

as many as the models may support. And

in some more complex models, we'll have

more than one parameter like lasso only

has the alpha. Um, technically it also

has this max iterations, but really the

only one that matters is this alpha.

Other models have many more

hyperparameters that we can um uh change

and so we want a way to systematically

test out those different combinations

and to see which one leads to the best

uh version of that model. Let's say the

best results. So um we're going to

explore that coming up. So we had lasso.

Um now this is where we ended last time.

We had ridge regression. If you guys

remember, this one is just a slightly

different penalty. Um,

it takes the it I drew it out for us. It

takes the same penalty we had before.

So, it has that um residual sum of

squares error, which is the main one we

use for linear regression, but it has a

penalty with an alpha. And then it has

the sum of the beta squares

beta i squares. So it penalizes it has a

penalty but it penalizes slightly

differently where it uses the square not

the absolute value. That's the ridge

regression. And this has the similar

effect of you don't in order to minimize

this right because our goal in training

a model was to minimize this thing

minimize this um quantity and find the

best betas that minimize this. Um so

generally yes you want to encourage

lower values but the um once you get

values that are a fraction if you square

them they actually get smaller. Um, so,

uh, it's it's not, um, it's not

necessary to shrink them all the way to

zero. They will get smaller as soon as

they're kind of below one. Um, so they

don't encourage it to completely go

away, uh, like the absolute value does.

It's just slightly different

minimization. Um, so what we see with

the ridge is we don't see the features

kind of get wiped out completely like we

do with a lasso. and lasso they get

encouraged to be um to become zero

because that's kind of the only way to

minimize an absolute value. But with

squares they can keep getting smaller

and smaller and smaller um fractions and

they don't have to become zero. It's not

as harsh of a of a penalty.

Um so uh the ridge was easy to use as

well. Um and it also has an alpha that

we can set. So, it's literally the same

exact code, just a different model, just

slightly different penalty, and it

results in different coefficients. You

notice that none of them are exactly

zero. Like with the lasso, you can get

ones that are exactly zero. We don't see

that with the ridge. You remember that.

Um, so we we s pointed out that last

time. Notice the coefficients aren't

zero. Um, and then we can evaluate it.

So we did our MSE calculation which is a

pretty standard thing where we use our

model to predict on a training set,

predict on a test set, evaluate those um

by computing the metric like the mean

squared error and we can see if we're

overfitting underfitting. This is

definitely the same kind of story we've

seen with all these models is

underfitting because the error is so big

across both sets

across training and test. So it's it's

definitely underfitting.

Um

and same thing as lasso, it has a cross

validation uh variation on it that

allows you to try out different alphas

and um do different folds. So 10 folds,

five folds, whatever, and compute the um

try to find the best alpha that way.

Okay.

All right.

Any questions on this so far from last

time from reviewing that a little bit?

Hopefully that uh hopefully that is

jogging your memory a little bit on

ridge and lasso. Um you know where we're

going to pick it up today is to finish

out this lesson with one more model

which is going to be a combination of

ridge and lasso. So you can actually

combine them together

um in a linear fashion those penalties.

So you can actually have both penalties,

the absolute value and the square. And

when you have both penalties um that's a

special model called the elastic net uh

regression or elastic net model. Um so

this is a combination of lasso and ridge

together. So you have lasso, you have

ridge and then you have elastic net

which combines both of those penalties.

Um let me show you the equation.

So here is the uh so here is the the

model. This is the same that we've

always had. This is our usual u model

fitting for linear. This is a basic

linear regression um loss function or

objective function that we're trying to

minimize to find the betas. Notice how

we have both of our penalties though

this time. So instead of just having one

of the penalties, we actually have both.

So we have the lasso penalty

and then we have the ridge penalty here.

So we actually use both of them and um

try to find a balance of minimizing

those two uh those two penalties.

Okay. And notice how they instead of

just a single alpha, we kind of have a

balance on both of them.

So, we can actually weight the lasso one

more. We can weight the ridge one more.

We can weight them the same. Uh we can

um change that around as much as we

want. So, they have two different

weights there um that they could be.

Um now what happens in reality is uh

we're going to see this in the model is

that um usually what happens is these

get combined into a fraction. So there's

usually a ratio of lambda 1 to lambda 2

and this is known as the um this is

sometimes known as the L1 ratio

and this is a this is a a parameter

inside the model that we'll be able to

set um along with alpha. So we'll be

able to set an alpha and then this

ratio. Um the idea is is that um the

ratio will uh allow us to control which

one is more dominant. So if this number

is bigger the um this lasso penalty will

will be weighted more. If this ratio is

smaller if it's less than one for

example that means that the um ridge

regression is more uh dominant. Um but

the so we'll have this we'll have really

this and this at our disposal and alpha

is um

alpha is kind of like a a you can think

of it as a scale that is um so lambda 1

kind of like lambda 1 plus lambda 2 um

combined to equal alpha.

So it's like our total level of penalty

um our total level of penalty and we can

set that equal to one. We can set it

equal to whatever we want. Um and so

these will be in this ratio and there'll

be a total level of penalty that we can

apply. So the model will actually use

these two parameters when we when we do

it. But that's how they're that's how

they're all related.

Okay. So ridge uses both penalties.

That's the only difference between lasso

or sorry elastic net uses both

penalties. Um so one thing I want you to

notice is that uh if we um if we want we

could set this L1 ratio all the way to

zero

um which uh if we do that um the only

way this L1 ratio could be zero would be

if lambda 1 is zero. So it would just

revert back to ridge regression. So it

complet if if this is zero this will

wipe out this term and we'll be back to

ridge if the L1 ratio is zero.

Okay.

All right. So we have a elastic net

model. Um now it's used the exact same

way as we did the other models in the

code. So we have elastic net um uh from

the scikitlearn linear model family just

exactly where we had linear regression

lasso ridge all of those came from this

linear model um elastic net also comes

from there and then the cross validation

version also comes from there um so

let's see so when we build our model

it's going to be um very very simple

easy stuff because it's the same code

that we always have um we just use the

elastic net. We set an alpha alpha

equals 1 is pretty standard um just like

it is in in the last one ridge that's

industry standard is one and then an L

L1 ratio of.5

that's pretty standard as well. What the

L1 ratio.5 is is kind of a um

uh kind of a that means that the lambda

1 to lambda 2 ratio is 1/2. Um, so

that's that's a pretty standard uh ratio

as well, but again, we could set this

equal to one and they'd be kind of

equally weighted. Um, L1 ratio of a half

means that the uh ridge regard the the

ridge penalty is a little bit more

weighted uh in that in that situation.

Okay.

So uh once we have this model um we can

do ffit and we can run that on our

training data and we can um get we can

figure out what our parameters are like

our coefficients and our intercepts. Our

model will have that but more

importantly we can use our model to

predict right so we can predict on the

test set. Um let me go back and load our

data and actually run this.

So, we're going to be using the same

data that we did for

uh lasso,

which is the I'm scrolling back up so I

can load it. It's the baseball data

here.

Um,

just run it from there.

It's this hitters.csv. So, hopefully you

have that one.

Let me load this.

Okay, so we loaded that and then that

should load.

Drop that unnamed column.

We will get our dummies

and then concatenate those split

scale. I'm just rerunning things. I'm

rerunning things so we can see our model

one more time.

So rerun that. Take a look at that. That

looks good. and then

fill in the NLES on the on those.

Okay. So, we should be able to run our

uh elastic net now.

Okay. So, let's import that and then

let's build our model. So, there we go.

We build our model and the intercept is

that. Now, of course, we can look at our

coefficients. Let's look at that.

Look at our coefficients. So remember

the coefficients are the uh betas. These

are our betas that are in our model. Um

so we can take a look at those. Now um

they're it's somewhere in between. It's

not a full lasso where we're going to

see some of these be zero. It's not a

full ridge. Um so the coefficients we

get are different. They're somewhere in

between there. Those two models that

we've already built. So not quite the

same um somewhere in between there.

Um and then we can use our model to make

predictions and and compute the MSE

uh or the RMSSE I should say as well. So

we can take the mean squared error, pass

that into the square root and compute

the RMSSE. So still pretty bad. Um this

is right around that 300 range of what

we've gotten for our other RMSSE. So,

it's not like elastic net is any better

than those other like linear or lasso or

ridge. And that's not surprising because

it's just adding those extra penalties.

We don't expect it to magically get

better. It's actually a more complex

um when we add when we add those in,

we're actually reducing it and making it

simpler. And we need something more

complex, I should say. So, we're making

it simpler um by by making penalizing

our weights a little bit more. And so,

it's still not a good fit. That's not

really surprising, right? It's still not

really a great fit.

And we can we can even double check

that. We know our RMSSE is pretty bad.

Um but we can double check it with this

R2 score. And it's, you know, still not

good. Remember, a one would be really

good. Um that'd be like a perfect linear

model. This is um still pretty bad.

Okay, so as we said, the alpha controls

the overall strength. Um so the higher

the alpha, the more overall penalty

we're supplying, which makes the model

simpler. Um uh but the L1 um ratio

determines the mix or that ratio of the

lambdas, the lasso to the ridge. Um if

you have it be um exactly zero, you you

revert all the way back to um if you if

you put it at zero, you revert all the

way back to ridge. One would be all the

way to pure lassos. Somewhere in

between, like one half is is good.

Okay,

so this is another example of trying out

different values of alpha in the CV to

see which one works. Now again, I'm

going to show us in a minute a

systematic way to do this, but this is

just trying out um different alphas that

we set up in this uh in this um

uh range. So we have different uh values

between minus2 and two um

logarithmically.

Um so these are uh logarithm values that

are between this between minus2 and two

and we choose a 100 different alphas and

then we choose a 100 different um L1

ratios between 0.01 and one and we run

that we run this um cross validation

with 10folds. So this is quite a bit.

So, we're doing 10 folds and we're

trying out a hundred different um

options. Uh every time we do an option,

we're trying out 10 folds to evaluate

it. So, it's going to take a minute to

run.

It's still running here. But again, what

this is doing is trying out different

alphas and it's it's going to do a cross

validation. And you guys remember the

t-fold cross validation is where we take

our data and we divide it into 10 folds

and then we um train on nine of those

and then test on the remaining fold and

then we rotate all the folds 10 times.

and that we average those mean squared

error metrics together um against those

10 different uh fold options to generate

a basically like an average performance

for that value of alpha. And we're doing

that a 100 times for all these different

100 alphas that there are and 100

different L1 ratios that we're trying

with them.

So that's quite a bit of processing but

uh it did finish.

So we can see what our best alpha is and

our best one ratio. So we get the best

alpha is this best one ratio is this. Um

and therefore we can uh build a model

with those with just these two guys as

the alpha and the L1 and um see how that

performs.

We build that model and then we predict

on the test set and we generate the

RMSSE. It's just a little bit better.

It's still not It's just a little bit

better, but it's still not good, right?

It's still 338. It is just way too big.

Remember, this is RMSSE, so it's in the

units of our uh target variable. So,

it's in the units of if we go back to

our data, actually, I could just print

it out here.

um this RMSSE.

If I just do this, we could take a look

at um DF

or I could look at Y test

and you can see some of these values.

These are these salary values in the

hundreds, right? Some of them are in the

thousands. Um but an error of like 338

is just too big. That's a really big

error. That means we would be off by an

average of 300 when our our values if we

just do the mean

um

is only 550 as on average is 550 but we

have this amount of error on average um

so that's just a way too big of a

proportion of error right it's not a

very good model and again we can verify

that by looking at this R2 for.

So if we go down here,

still not very good.

Here's some of our coefficients. So

remember, you can always take your

coefficients and line them up to your

your data columns. Uh so that you can

get a sense of what coefficient belongs

with what feature. So that's all we're

doing here is just creating a series

where those coefficients instead of just

printing out the coefficients, we're

actually lining them up to the columns.

So this tells us um remember the larger

it is the more influence it kind of has

on the on the final result. Um either

way, so like this has a big negative

influence. um this has a large positive

influence.

Okay, let me pause there. Any questions

about the

elastic net model?

This is a really this model is a really

good one to use when you are building a

linear regression and it's performing

well but it's overfitting. This is a

really good one to use because you can

balance

lasso and ridge you can get the best of

both worlds. So the the main strategy is

if you are using a linear regression and

you see overfitting

um meaning that it's performing decently

so on the training set

it's performing okay but then on the

test set like you know it's it's not

underfitting. it's performing pretty

well on the training set, but then on

the test set it's um performance is much

worse. That's overfitting. If you're

overfitting, then this is a great model

to use because we can try basically by

by rotating through different alphas and

different L1 ratios, we can try out

different strengths of penalty and

different variations on lasso and ridge

together. This is a really good model to

to use for those overfitting cases where

linear regression is doing decently. Um,

but it's overfitting,

right? So far, we haven't ran that case

because so far, no matter what model

we've used, it's always underfit. So,

anytime we have those underfitting

cases, it signals that we should likely

just use a more complex model. And we

haven't learned about those yet.

um we will coming up in lesson four, but

um that's for this data. That's ultim

ultimately what we'd want to do is

probably use a more advanced model

because it's underfitting um just using

a linear regression and and then using

the the overfitting variations of linear

regression like lasso ridge and elastic

net.

Okay.

Any questions on this on elastic then

the TV? Yeah, we Yeah, I think I have

it. I can share it with you.

I said that and now I can't find it. I

thought I had it.

I don't have it. I thought I had it, but

I don't.

If anyone does have that one.

Yeah, I'll look one more time. I thought

I had that one.

Um,

yeah, it's not in there. I had it. Let

me see.

Yeah, I don't have it either. I thought

I had it in here.

Yeah, I don't have that one.

I don't have that one. I'll have to find

it. Uh I have this marketing data. I

don't think this is the same one.

I have this marketing data. I don't

think that's the right one, but you can

take a look at it.

No, we're using So, for this example,

we're using the same hitters data set

that we used earlier for lasso.

No, that's an earlier one.

That's from the uh very beginning of the

notebook. So that's the that's from this

one.

Oh, this Oh, this is where it is. Sorry.

This is where it is. You can find it

here.

That's right. It was from a URL.

It was used in the very beginning of the

notebook.

And we did we did this.

Okay.

That's right. That's why I didn't have

it downloaded.

Okay.

All right. Any other questions on the

elastic before I move I'm going to move

on to uh finding those a systematic way

to find the best hyperparameters.

Um, I'm going to show you a couple

strategies to doing that. Um, so far

we've just ran CV with some random

choices. Um, I'm going to show you a

better, more systematic approach. That's

kind of the industry standard for doing

tuning. Um, so I'm going to I'm going to

show you that next, but any questions on

the elastic net?

Okay. And again like you know

scikitlearn makes it really easy for you

guys

because it just behaves the same way as

any other model. You use the object and

then you do ffit and predict right? So

the ffit is going to train it um and the

predict is going to allow you to use

that model to predict. It's it's super

easy that way. Every scikitlearn model

is like that dofitit and predict. So it

provides a really simple way to use

basically every model.

Okay,

let's talk about let's finish up this

lesson with a couple things. Um, one of

those things is going to be

hyperparameter tuning. So what is this?

The hyperparameter tuning is a

systematic way to find the best

parameters in a machine learning model.

So a lot of machine learning models have

what are called hyperparameters.

These are not the betas that we learn

during the training that's learned from

the data. These are settings that we set

ahead of time like the alpha. That's a

perfect example like alpha L1 ratio in

in the elastic net. We set those up

ahead of time and depending on what we

pick for those we get different

performance, right? And so what we

really need is a systematic way to find

the best settings for those

hyperparameters as we are training our

models. Um the the the main like idea

behind this process though is going to

be to systematically try out different

combinations as many as we want to try.

And so we're we're basically going to

have a strategy for tuning that is going

to exhaust all the combinations of those

hyperparameters that we want to try

until we find the one that performs the

best. Um and and that strategy is known

as grid search. Um and essentially what

it does is it sets up a grid um where

which is basically like a matrix to say

okay which parameters do you want to

try? I want to try um alpha and I want

to try L1 ratio

um L1 ratio like let's say I want to try

these two. So we set these up in a grid

where we say, "Okay, I want to try this

value. I want to try this value. I want

to try this value. This one, this one,

this one, and on and as many as we want

to try." So we could set up set those up

systematically like a linear um a

linearly spaced like I want to try every

alpha between between 0 and 10 spaced by

one um whatever. You know, we can set up

different ranges of those, but that's

going to be in this grid. And then the

L1 ratio, same thing. We can try out

different values of these that we want

to try. Let's say there's many of those.

Um maybe every um tenth between 0 to one

I want to try out. Um so you set up your

parameters and you can set up as many as

you want in the grid. And then

essentially what you're going to do to

do grid search is you're going to work

your way through every combination of

those. So you're going to try out this

combo. You're going to try out this

combo. You're going to try out this

combo.

this combo. So the first value of alpha

with every possible L1 ratio, then go to

the next, try out the next value of

alpha with every L1 ratio, and on and on

and on. So we're going to try

all combos

in the grid.

We're going to try all combos and we're

going to find the lowest MSE

combination. find lowest

MSE

combo.

So whatever leads to the best model is

going to be the um parameters that are

that are deemed to be the best. And the

idea is once we have found those we know

that we can use we can go ahead and

train a model with those best alpha and

len ratio and on and on and on.

Yeah, when you get an So this goes back

to the error. Remember that for a

regression,

the error is this measurement of how far

off we are, right? So if we have a bunch

of points and we draw we fit a line

through there, the the MSE is measuring

this distance, right? So what do you

think is a good distance? Like if our

model is perfect,

what's the best distance from our

predictions to the actual points? Zero.

Yes. So the lower the better. The lower

the better. Um so for an R RMSSE, the

lower the closer to zero the better.

However, the RMSSE can be it's its units

are interpreted in the units of our

target.

So what is deemed to be good is relative

to our target. Like let's say our target

is in the thousands. Like it averages in

the thousands. If we produce an MSE of

50 or sorry an RMSSE of 50, that's

pretty good, right? Because our units

are in the thousands

and we're only on average we are off by

50 units,

right? Our distance away is about 50

units. That's pretty good. So the RMSSE

is relative to your target variable.

Does that make sense? Yeah. It depends

on the target. It depends on what you're

trying to predict.

So that's why we got RMSSE that were in

the 300s for those hitters, but the

average was the average of the target

was in the 500s. So that's a really bad

proportion of error relative to the

average target value. Right? If our

RMSSE was 300,

but the target was sitting in the 500s,

that's just too much error. Way too much

error, right? That's just too big of a

value. Um, our predictions are just off

way too much

in terms of that distance. So, this

would be this was a bad model. It was

underfit.

We know that from the the RMSSE. So

yeah, the RMSSE closer to zero, no

matter what is good,

zero is being perfect. Um, but it to

know what's good, you need to know what

your target is on average and then think

of this as kind of a ratio to that

average target. I think that's the best

way to think about it.

Okay, so going back to this grid idea is

so the grid is just basically laying out

all possible parameter combinations and

trying them all out by fitting and

predicting until and generating an a

metric like an MSE

until we find the one with the lowest

MSE. So find the lowest MSE combination

and that will be the best

that will be the best combo and then if

we once we know that best combo we can

use that we can use that alpha we can

use that L1 ratio and use that model

going forward we can we can use those

parameters in our model so this strategy

it has a name it's known as grid search

so it is a hyperparameter tuning process

that tries out all combinations S.

So what's the what's the U benefit to

this is that we get to test out a lot of

different combo combos of those

parameters like the alpha and L1. So we

can be confident what the best model is,

right? So we can pick the alpha and L1

perfectly because we're trying out a

bunch of different combinations on the

data to see which one's the best. What's

the downside?

It's expensive, right? It's an

exhaustive search. So if you have many

different parameters and you're trying

out many different combinations, it can

get exponentially

expensive

to perform this search. Okay, so grid

search is great except for the fact that

it can be expensive if you have many

parameters with with very wide ranges

that you're searching over because that

that's a lot of combinations you have to

test, right? And especially if you have

a lot of data, that's going to be

expensive

um to do.

So, uh we're going to practice doing

grid search, but that is that's the pro

and con. The pro is that we get to try

out all these combinations and see which

one's the best. The downside is it can

be expensive to do that if you have a

lot of parameters um that you want to

tune for your model um and you have very

uh many different choices that you're

trying to evaluate for those and it just

creates a really big um collection of

combinations that you have to try out,

right? Um that's the only downside to

grid search.

Now on the opposite end of the spectrum

of that is a randomized search or random

search and this will basically just um

do a sampling of those parameters from

um kind of fixed uh specified

distribution. So essentially what you do

is similarly you define your range. So

you say I want to look at alphas um

between zero sorry between let's say

yeah 0 to 10. I want to look at a bunch

of different alphas. Um, and I want to

look at a bunch of different L1 ratios

that are between 0ero to one.

0 to one. And um, what we do is we say,

okay, I'm going to restrict only testing

20, 30, 40 times. I'm not going to do

all possible combinations. I'm just

going to randomly sample something in

this range and randomly sample something

in this range. And so and I'm going to

perform that experiment a fixed number

of times. So let's say I set the uh

sampling where I'm only going to do um

20 evaluations.

And so 20 times we're going to pick a

combo randomly. So I'm going to pick an

alpha and I'm going to pick an L1 ratio.

L1 ratio.

And um we are we are just going to uh

sample those randomly from this range.

Um and we're going to use those and test

those out and then it's but otherwise

it's the same as grid search. Whatever

is the lowest MSE

um so whatever is the lowest MSE is the

best.

So we evaluate those. We sample we train

the model. Evaluate it. Whatever is the

lowest MSE

is the best is the best combo. Now,

what's the benefit to this is it's a

much more controlled experiment in the

sense that we um aren't going to iterate

through every possible combination in

the grid. We're we basically set up a

fixed number of times we're going to try

out stuff.

The risk to doing this is that you're

not you're not exploring all

combinations, right? Because you're

randomly sampling, you may get unlucky

and you may not stumble into the best.

You you can make um samples and figure

out what's the best amongst your

samples, but you may not be covering all

the combinations. Does that make sense?

The grid search is going to try every

combo. The random search is going to

randomly sample those combos.

So, it's not going to try every single

one. It's going to try a limited number,

however many you set up. Now, if you set

that number really, really, really high.

Now, you're starting to approach a grid

search because now you're sampling so

many of those combos that you basically

are trying them all at that point,

right? Um,

so, so that's the way the random search.

So by the way, both of these use cross

validation in the sense that when you

evaluate accommodation, you're actually

doing it with cross validation. So when

you do an evaluation, you're going to do

probably 10 or five folds where you

split your data and then you test it on

the rest of the folds and evaluate or

train it on the rest of the folds,

evaluate it on one of them and generate

an average MSE to get your evaluation.

So every evaluation is using cross

validation.

That's why that's and hopefully you can

see why this would be so expensive for a

really big grid, right? Because you're

trying out many different combinations

and every combination is going to do a

cross validation procedure. So it's

going to train 10 times and test against

10 different folds and average those

together. it's going to be a pretty

expensive operation

for a really big grid, right, of of

parameters.

Um, but these are the two kind of

systematic approaches we have at trying

out different hyperparameters. Remember

those those things are called

hyperparameters. These are those choices

that we have before we train our model.

Um, those choices we have that affect

the performance of the model like the

alphas, the L1 ratios, those kind of

things. um we have control over what

they're going to be. This is a

systematic approach to find out what the

best

uh value of those parameters is going to

be on our data.

Right?

Okay. So before we practice this, we're

going to practice with grid search

first. Um

any questions?

Uh, I don't know if it has a built-in

That's a good question. By time limit, I

don't know if it has a built-in way of

doing it, but you could certainly set up

like a a a loop um to like to wrap

around. Do you know what I mean? Like

you could set up a loop where you check

the time if it's if if the time elapsed

as you're doing a search if the time

elapsed is greater than the the time

limit then you can kind of break early.

Um so it's not hard to implement that

but I don't know if it has that built

in. I don't think it does

because I don't think it really cares

how long every evaluation takes. It's

just going to exhaust all those

especially in a grid search.

But um yeah, I there's probably a way to

manually kind of set up a time time

loop.

So hyperparameters are um settings that

we have on the model itself. And a

really good example of this is like the

alpha and L1 ratio in the in the elastic

net. So they're not things that we um

learn from the data directly like the

betas in the model like those get

trained directly by doing the um least

squares process right um by doing that

gradient descent and all that

optimization.

Um so these are not learned from that.

They're actually set ahead of time. And

so what we're saying is the best way to

understand the effects of those is to

try out different combinations of those

until we land on the best one. Right? So

hyperparameters are those options we

have in the model like the alpha like

the alpha and l1 ratio in the uh elastic

net. Many models have hyperparameters.

Um we're actually going to see that in

in future models that we study. they

have options that you can set that

affect their performance.

And so this this is just a strategy to

evaluate those different options to see

which one's the best.

Yeah. So again, hyperparameters, those

are settings on the model itself um that

affect the performance of it.

And basically we have the two two

strategies here. We can set up an

exhaustive grid and search through all

of those until we find the lowest MSE uh

option or we can randomly sample

potential options, try them out and see

which one's the lowest as well. And do

that a fixed number of times. Um

sort of like a fixed number of trials

almost. um which has a risk of not

trying out every option but but

hopefully you try out enough that you've

explored the space a bit and you get

some quality choices there but no

guarantees right no guarantees you try

everything which is what a grid search

will do it will try everything

okay now luckily per usual scikitlearn

has something to manage this process for

us in terms of grid search. Um so in

that way we will not need to manage this

process ourselves. We can just rely on

scikitlearn. And so if you're doing

hyperparameter tuning um this is going

to come from the model selection module

inside of sklearn. So we're going to

import from from sklearn the model

selection module. We have our grid

search cross validation.

Okay, that's what the CV stands for.

grid search cross validation. So, this

is going to do that grid search

strategy. Um, we're going to set it up

with our dictionary essentially of

choices. So, we're going to say, hey,

here's the alphas I want to try. Here's

the L1 ratios I want to try. Um, and

here's my other settings like uh how

many folds I want to use, what my random

state is for the shuffling. So, we'll

set all that up. Um,

and then we'll just run the grid search.

And then what should come out of that is

the best options for our parameters from

the grid and then we can use those going

forward in the we can build a model with

those best options right so we're really

doing some evaluation here of what is

going to be those best alphas those best

0 to1 ratios on our data set right and

the only way to really know that is to

evaluate them because they're not things

that are learned during the training

hopefully that makes sense right they're

not things that we learn directly from

training. There are things that we have

to set and then kind of evaluate and see

how they affect things.

Okay, so we have grid search CV. That's

going to be our primary um tool to do

the evaluations of the different

hyperparameter options.

Grid search CV. Um we're going to set up

our cross validation uh object here. Now

I want you to pay attention to this is

that um it's a slightly different

version than the kfold we had earlier.

So we've used k-fold before with a

certain number of folds. This would be

10 folds and we can set a random state

for the shuffling um that happens in the

folds.

But this is actually a slight different

variation on it where it is a repeated

kfold where we do three repeated trials.

Now why would we do that? It's to be

extra extra extra careful with the

shuffling.

So this what this means is we do three

different shuffles. So we do kfold, we

actually repeat it three times with

three different shufflings. That's all

that means. So the repeated kfold is

actually a bit beyond the just basic

kfold. What basic kfold will do will

we'll will shuffle and then do our

splits into 10 splits and then train on

nine of those. test on the other one and

rotate through all the splits.

We're actually going to do that process

three different times with three

different shuffles. So this and we're

going to average 30 results instead of

just 10. So repeated kfold is just going

above and beyond to do extra to repeat

the kfold three different times. In this

case only three. We could do more,

but um now is that necessary to do? You

could argue not necessarily. Um but it

just provides extra robustness

uh beyond just our single shuffle and

then split and then rotation of those

folds, right? We're doing it actually

three different shuffles. Um so we're

repeating our kfold three times uh for

every now is the thing is we're doing

that for every evaluation. So it is

going to be more expensive than just a

basic K-fold.

So we have three different K-fold trials

that we're doing essentially.

Okay, hopefully that makes sense. This

is the repeated K-fold. We haven't

really seen that before. we've only

worked with the Kfold, which would get

rid of this repeats option and only have

uh 10 splits in a random state for the

for the single shuffle that we do. So,

we can um recreate that same shuffle

every time. Um but now we're actually

going to do three random shuffles, uh

three different trials. So, one shuffle

creates the and then create the 10

splits, evaluate, then go back and do

another shuffle, another new 10 splits.

So, one thing that should be um clear is

that we get different splits every time

because we're going to shuffle once,

right? We're going to shuffle once and

generate our splits

and then we're going to shuffle again,

generate these splits which are going to

be different and then shuffle one more

time for for three different times,

right? And then get get these splits and

then we're going to get 10 metrics here,

10 metrics here, 10 metrics here, and

then average all of those together.

So, it's a bit more just going up extra

above and beyond for a K-fold. Okay.

All right. So, here comes the fun of

when you do grid search. Now, the grid

is actually just a dictionary. It's a

Python dictionary where you declare what

your parameters are going to be inside

the dictionary and you set up a range of

values that you're go or a list. It can

be a list. It can be a range

but some declaration of what you are

going to test and evaluate inside of

your grid search. So the grid is

initialized as an empty dictionary.

And then what we do is we say okay in my

grid I want to check different alphas.

So we're going to add a collection of

alphas in here that we're going to test.

So let me make a comment there. We add a

add a range of alphas to test. And this

range is a this is just like the Python

range. Um

this is just like a Python range um uh

operator here where this is going to be

uh every so it's going to be um every

uh value between

zero and one um uh steps with a step

size

of 0.1. So, it's going to try a bunch of

different alphas um between uh zero and

0.1

sorry 0 and one stepping by 0.1. So,

it's going to try zero.1

2.3 point 4.5 6 right all the way up to

one.

So, that's what this will do. And it's a

numpy range. So, it's just all those

decimals between 0 to one.

It you can use either that's valid.

Yeah, you can do you can do that to

create a dictionary or you can use the

keyword um dict. You can use either one.

Either one works.

Whatever whatever you want to use.

They're the same.

Yeah. The the reason people prefer

dictionary is because um sets are

created with the same braces.

So it it makes it clear what you're

creating as a dictionary. If you use if

you use this, that's the only advantage

is it's just plainly obvious what you're

making. Uh because technically you can

make a set with the curly braces as

well.

Yeah,

no worries. Um okay, so we have our

alphas here. So what I want you to

notice is that we are going to try out

different alphas and we are that's the

only parameter we are going to try in

our ridge regression. So we're going to

we're going to try ridge but just try

different alphas in the in this range um

in our grid search. So the grid search

CV takes in a model. It takes in our

grid dictionary which is really

critical. We need that dictionary to

declare what we're going to try.

um we need a scoring to say to find the

best. Now remember it uses the negative

to find the lowest which is going to be

the the least negative option.

Um otherwise it wouldn't um just based

on the optimization it would look for

the highest value. Um so the highest

would be closest to zero in this

situation. Um so we use negative and

again we could use squared error. It's

using absolute. We could use um squared

uh either either one works.

Um more typical would probably be

squared error, but um absolute is fine.

Here's where we have our repeated kfold.

So we pass in our um how we're doing CV.

That can be it can be a kfold object. It

can actually just be an integer, which

is say I just want to do 10 splits or

five splits um to to do every

evaluation. But these are the bare

minimum that you need. Just really the

model and the grid and your CV. Um what

metric you're using to evaluate what's

going to be the best. And then this end

jobs is to parallelize. If you have it

set to minus one, it's going to it's

going to try out all the grid options in

parallel. Um which is nice. It's going

to help speed up the overall search.

Okay. So let me mark that down as n

jobs equals minus one.

tries out the combos in parallel.

So in this situation, we actually don't

have more than one parameter. We only

have the alpha. So we're really just

going to be systematically working our

way through every alpha and evaluating

which one's the best right with this.

And notice that in order to use this

grid search, all we have to do is call

search.fit. So it works kind of like

every other model does, right? It's the

grid search.fit.

And we pass in our data.

And we um once we're once this prints

out the results, you get a results

object um which has a best score and

then a dictionary with your best

parameters. So, whatever your best grid

member was or grid members, um it prints

that out and you can So, for from that,

we can um grab our best alpha, which

which let's confirm what that ends up

being.

Oops. We need to import repeated kfold.

So we'll import that.

Oh, I didn't. Let's do from

sklearn.linear

model import ridge.

Okay.

Okay. So, it completed the search and

what we found is this is the best score

is 238 for the mean absolute error and

the best alpha that we got was 0.9. So,

the best alpha that worked here, the one

that gave us the best score was actually

0.9 as the alpha. So what it did is it

tried out everything between this range

and 0.9 was the best. So it did cross

validation, tried out every single combo

in our grid.

So if we want we could actually print

out

print our grid so we can see

what our combinations were.

So, it tried out all of these guys and

the best one that we had was 0.9.

Okay, so pretty cool how that works. And

if we had other parameters, like if we

were doing a elastic net, we could add

those into our dictionary and it would

do all combinations of those. So if we

did um so for instance to add to our

grid we could do grid

um L1 ratio

this would be for like an elastic net

right now the ridge regression by itself

doesn't have an L1 ratio parameter but

just as an example um we could try out

different ranges um similar range

different one um maybe an exact list

whatever we want to do. So this is going

to try out different ones between 0ero

to one

as well. And so it's going to try out

every combination of these from this

grid.

Okay, if we did that. But again, this

the ridge regression doesn't have an L1

ratio. The elastic net does. So that the

ridge regression only has an alpha to as

a hyperparameter. So we're only testing

out that one.

Okay. So that's grid search CV.

Pretty useful. This is pretty useful in

doing parameter tuning again when you

want to try out ranges of different

values and you can evaluate those to see

which one is your best and then we can

use that best going forward. So we can

for instance this is what this code does

below it is it fetches the best. Um you

can do it this way or you can do it um

the alternative is to do results.b best

params

and then you can just grab it like this

alpha.

Either way you can do get or like this

um and it this is just a dictionary,

right? And you can grab your alpha. So

that's the 0.9 um and we can pass that

alpha into the ridge regression and go

back and refit it to our data um and

then use that model going forward. So

the grid search really just evaluates

those different options, tells you

what's the best according to this score,

right?

And you should, by the way, you should

interpret this score in the positive

sense. It's only negative because we're

purposely making it negative to find out

what the lowest option is, right?

Because the lower is the better. So we

we purposely make it negative to make it

whatever is the least negative is the

winner. Um more negative is worse.

So it's really positive version of it is

the is the true result for the error. Um

and they are a tool from scikitlearn to

put together your model with your

pre-processing steps. So they kind of

get automated together. Um and they

combine everything into kind of a

streamline process. You're going to see

what that looks like, but it's a really

nice um feature of scikitlearn. Um why

would we care about pipelines? They help

organize our code um so that we ensure

that we basically always run the

pre-processing steps before we train and

use a model to with the predictions. Um,

so it bundles those steps together,

minimizes the risk of forgetting a step

because one of the things that can

happen is when you do pre-processing, if

you're doing it on the training set, you

have to do it on new test data as well

when you put it through your model

because your model is training against

that pre-processed data.

So in order to make sure you never

forget that, you can bundle it all

together in a pipeline which is going to

make things really really easy to use

and and make sure that those steps

happen in a repeatable way. Um and it

makes things easier to uh deploy that

model as well because everything is

together in one pipeline. So in the in

the industry, I've seen this a lot. Um

you know, people will do their initial

exploration steps and initial model

building. They may not use pipelines

right away, but as they found their

model, um they'll generally move it into

a pipeline and all their steps into a

pipeline so that it's uh easier to work

with um when you're when you're

deploying it, actually using it uh in in

the real world. Um, so here's what a

pipeline generally looks like. It's from

scikitlearn. It's this pipeline object.

Um, and the pipeline is made up of steps

that we're going to see that that are

various um uh basically um kinds of

pre-processing we've seen before like a

scaler or um filling in missing values.

Those kind of things we can put here in

the steps which is basically a list. um

steps is just going to be a list of

scikitlearn functions that we can apply

to data. One of those being a model. Um

and then whenever we use the pipeline,

it's basically um you know, it's going

to be something like pipeline.fit

or pipeline.predict.

So the pipeline kind of behaves like a

model. It's just going to contain many

more steps than that like the

pre-processing steps we've worked with

before. Um, and it also has some

capabilities for caching. So you can

like uh cache some of the data in

memory. Um, so that if you're reusing

the predictions, it kind of goes faster.

Um, so there's some options for that

too. I'm not too concerned about that at

this stage, but the main thing is going

to be filling out our steps and then

using the pipeline.

Okay.

Um, so some important bits of

information about the pipeline is that

it is going to be a sequence of data

transformations that will have at the

very end of the pipeline the model

because of course we're going to do

transformations and then train a model

or predict with a model. So every

Oh, can you guys hear me? Okay,

not able to hear me. Thanks for letting

me know. Can you guys were you able to

hear me so far?

Okay. Make sure. Yeah, it might be on

your internet or your your uh Yeah, it

seems like seems like it's good. So,

no, you can't hear me. Check your

volume. Check your headphones if you're

wearing headphones. Oh, no issues. Okay,

perfect.

Okay. Yeah, local internet issue. Yeah.

Okay.

Always let me know. always let me know

cuz it could be the case that it is me.

So, um always always make sure to let me

know. Um but sounds like yeah, you may

want to check on that. Um

so, okay. What I was saying is every

pipeline's going to have a uh a sequence

of steps that go first and then the

model at the end. Um so, the order

really matters. Um

uh so the order matters in the sense

that we want our transformations to go

first. Things like scaling, things like

filling in missing values, we want those

to be first and then we want our uh

model to be last because we want those

transformations to happen prior to

training or prior to prediction. So

usually what you'll see in these

pipelines is a model at the end, right?

a model that's going to be at the end of

the pipeline because we want basically

our processing steps then our training

or our processing steps then our

predictions. Um so everything in the

pipeline though is going to be from

scikitlearn. Uh that's how it gets

automated in the sense that all of those

things are going to have fit and

transform functions built into them so

the pipeline can use them. Uh, and then

the last step is going to be a model

that has a fit and a predict. So it's

pretty standard that the last part of

the pipeline is just going to be a

model. Um,

uh,

so we can um, as we do more modeling,

we're going to play around with the

pipelines quite a bit and see how we can

change up some of the parameters. like

if we want to change a model's parameter

um we can actually adjust it to do

things like uh grid search or cross

validation. So um we're going to see

some examples of some pipelines but for

right now mostly what we're going to see

is how to build one and then how to use

one. And then as we get into lesson

four, we'll get some more practice with

pipelines cuz we're going to start using

them quite a bit uh to build our models

rather than do manual steps uh all the

manual pre-processing

um and then kind of building a model

from there. We'll just include all of it

together in a pipeline.

Okay, so the example we're going to do

is with this housing with ocean

proximity. So we've actually looked at

this data set before. Um so we have uh

this ocean proximity data set that has

the feature of like how close it is to

the ocean like the bay or the less than

1 hour. Remember we had that and it had

the median house value for different

neighborhoods. Um so we're going to work

with that one again. Let me make sure I

have that one uploaded.

You guys should have this one. It should

be in your uh data sets.

Um, I'll I can upload it here in case

you don't have it though.

Does this use multi-threading? I think

it does. Yeah, I think in order to do it

can do uh um I think it can do

processing in parallel for some of the

pipeline steps. Um, now does it use that

all the time? Not necessarily because

some of it is sequential in nature where

you have to do one step and then you do

the next step and then you do the next

step. So it's not like you can do them

in parallel.

Um in terms of the like you need to know

the output of one step to compute the

the output of the next step. Um so it

can but it it doesn't always lend itself

well. The thing that will use

multi-threading is is like the training

process could be parallelized

like the fit um can be for some models

it can be parallelized not every model

it so long answer is or the short answer

is that it depends

depends on what kind of transforms

you're doing and what kind of model

you're using if you can really take

advantage of

Okay. So, we load our data here and take

a look at that. Um, do you guys have

this data set? Are you able to load it

in? If you're following along, are you

able to load it?

Okay.

And and again, we've worked with this

data before, so hopefully it's somewhat

familiar. Remember, every row represents

a neighborhood and it has a we're going

to end up trying to predict this median

house value as our target um variable,

our dependent variable. Um and we're

going to use the rest of these features.

Remember that um this feature is in

particular going to need to be one hot

encoded,

right? It's going to be one hot encoded

because it is currently a string and we

need to turn that into a numerical

feature which is the one hot encoded

feature. So we're going to have to do

that but we're going to do that as part

of our pipeline.

Okay. So we'll be able to include that

in our pipeline steps uh to to do one

hot encoding which is nice.

All right. So we're going to split apart

our data um as we normally do. So we're

going to uh create our feature uh data

frame which is everything but this

median house value. So we go ahead and

drop that column and then our target is

the median house value. So it is just

that column here. Pretty standard. Um

and then we're going to train test split

and um split it into 30%

uh test data. And again random state you

can choose whatever you want to be. that

just affects the shuffling. Um, so

whatever doesn't really matter what it

is. It's just so that when you rerun

this, you get the same result in the in

the shuffle.

Okay, so we have our train and our test.

So you want to make sure you run those.

All right, so what we're going to do is

take a look at our data

and see if we have any null values. Um

if you guys remember this data actually

did have null values. You can see it

here in this this guy and exactly how

many there are is from this the sum. So

we have um 162 nles in in this data. Uh

and this is just a training data. So of

course you know the test data could have

that in there as well. Um so that's

something we're going to want to make

sure we fill in the blanks on any data

set we use whether we're using the

training or test set. Um, like if we're

doing training, we want to make sure

that gets filled in. If we're doing

predictions with the test set, want to

make sure that gets filled in. Um, so we

we should be doing that. Um, now

what we're going to do is use this data

to help uh train our pipeline or or use

with our pipeline. We need to construct

our pipeline. So far, we've just split

apart our data. We haven't done anything

with our processing steps in our model

yet. Um so roughly

it this should be the flow of our

pipeline. What should happen is we

should be doing some type of feature

scaling

um some type of uh feature um

manipulation. So that could be

engineering, that could be um that could

be uh doing the one hot encoding. Um so

extracting new features like one hot

encoding,

one hot encoding. Um we are going to be

doing that and and by the way, this is

split up into this is when we use our

pipeline for training.

Um it's going to look like this where we

do our scaling, we do one hot encoding,

um we have our model here. Um, so that

could be a linear regression, that could

be a lasso, that could be a ridge, it

could be elastic net. Whatever model we

end up using is going to be last in the

pipeline. And we're going to run this

pipeline. Ultimately, we're going to run

pipeline.fit,

right? We're going to run a fit function

and we get a fitted model as the result

of this pipeline.

Then when we use it when we use our

model for prediction,

we use our model for prediction in this

lower part. It's the same pipeline, same

exact pipeline, but it's this model has

now been trained.

So we now have a trained model here. So

the great thing about the pipeline is

it's the same this is the same pipeline

that we're using here. So it's just

going to it's going to repeat those same

transformations. It's going to do our

scaling. It's going to do our one hot

encoding. It's going to use our model

and it's going to generate predictions

and generate uh we can we can do

predictions. We can do evaluation like

in a cross validation. Um we can use it

however we want to use it. Uh but notice

that the pipeline makes it consistent

between training and test. We're using

the exact same transformations

and the model is last. It's it's either

being trained or it's being used for

prediction, but it's last. Our

transformations are upfront, which are

things like our scaling, things like our

one hot encoding, right? Those happen

first. No matter what data we put

through there, we put our training data

through there, we put our test data

through there, they're going to go

through the same steps,

right?

So that's that's the design of the

pipeline. That's what it's supposed to

do. So our job is to create those steps.

So we need to create those relevant

steps and then put them together into

this pipeline. Okay. So that's going to

be the code we're going to see coming up

is we're going to build out these steps

and then put them together into the

pipeline.

Um any questions on this diagram? Does

it make sense what we're trying to do

with this pipeline? We want to have

repeatable steps during the training,

during a prediction process.

Okay.

All right.

All right. So, um, a couple of things

we're going to need is, uh, to first of

all, let's jot down what steps we're

actually going to do. We're going to

need to deal with missing values. So,

we're going to fill in we're going to

need a pre-processing pre-processing

step that fills in any nulls. We always

need that, right? So, if there's nles,

we're going to fill them in somehow.

We're going to define how we do that in

our in our step. Um and we also need to

one hot encode and we need to scale

right those are pretty standard steps

that we've dealt with whenever we're

building these models right so pretty

standard things fill in nles one hot

encode any categorical data whatever

however much we have and then go ahead

and um standardize which is the scaling

so this this just is the same word for

scaling our numeric features so we're

going to we're going to define Windows.

Um, so that's why we're going to go

ahead and import from pre-processing.

We're going to import our scaler. Um,

again, we could use minmax scaler here.

We're going to use standard scaler. Um,

but we could use minmax. Um, we have our

one hot encoder here. Now, usually when

we do oneh hot encoding, we use pd.get

dummies. This does the same thing as

that, but because we're going to be

building a pipeline, we actually want

the scikitlearn version of git dummies.

So this is the scikitlearn version of

git dummies here. And it and we have to

use that version in the pipeline because

everything in the pipeline needs to be

an sklearn object. It needs to be an

sklearn tool or object.

So um instead of using pandis get

dummies we're using one hot encoder

which is does the same thing. Okay. In

fact it just this basically just uses

pd.get dummies um under the hood.

Okay. So it just uses that uh anyways.

It's just code that builds on builds on

that.

Now what's really nice here is we're

also going to use from sklearn.impute

impute. We're going to use a simple

imper now what this is is an automated

way to fill in missing values. So this

is a fancy way of basically doing the

the fill na on a data frame. So simple

imputer um we are going to basically

fill in the blanks. What we're going to

do when we create this object is give it

a strategy of how to fill in blanks.

Should you use the average? Should you

use the median? Should you use the max?

Should you use the min? should use a

default value. We're going to tell it

what to do in this object.

Okay. So, we're going to we're so we're

going to use this as our automated tool

for filling in missing values. So,

that's really nice. It has so this is

going to be a critical part of our

pipeline an imputer that's going to fill

in missing values.

So, we have that.

Yeah. Coding to reduce coding. Exactly.

Uh we have our pipeline now. So we have

our pipeline. So our pipeline is going

to hold everything. So we need the

pipeline object um to hold everything

and that comes from sklearn.pipeline.

Um so everything's going to actually go

into a pipeline object. We're going to

see how that looks. Um and finally we're

going to from skarn.compose we're going

to use a column transformer. The reason

we're going to do this is because we are

going to specify for some columns like

the numerical features we should be

scaling

for some columns like the categorical

features we should be one hot encoding.

So the column transformer will allow us

to map different transformations to

different sections of columns which is

really useful. So this is actually going

to be a critical part of our pipeline to

apply to make sure we only apply this to

numerical features and only apply this

to categorical features. Right? So this

column transformer will help us um to to

apply pre-processing to particular

columns. Um like that ocean proximity is

the only one that really needs this but

every other column is going to need this

all the numerical features.

So, we're going to use this column

transformer. And again, we're going to

see how this looks, but just trying to

give you an idea of why we're importing

all these things.

Okay. So, let's import those.

Uh, this mentions about the column

transformer. We just talked about it. It

allows us to have a particular column or

group of columns get the right

transformation. So again, uh, looking

ahead to our pipeline, the numerical

features are the ones that are going to

need scaling, but the categorical

features are the ones that are going to

need one hot encoding. However many

categoricals there are. In this case,

there's really only one, which is that

ocean proximity. Go back to our data.

Um, you can even see that in the info,

there's just that one. Um, and we see

that here, right? Just this one string

column that should be one hot encoded.

All these other guys should be scaled.

Right? They should all be uh uh standard

scaled.

So this will allow us to specify those

distinctions.

All right. So let's get started building

our pipeline. So this is going to be

really cool. We're going to build out

the pipeline. Um let's extract our

numerical data and our categorical data.

Now this is a really neat way of doing

that that I'm not sure we've seen

before.

Um so what this does is we'll take our

data frame particular our training data

frame and select our data

that's what this select dtypes does is

select data from it um which includes

only the object type columns so only the

object types. Now what's that?

The object type is the string right? So

this should select only this column

because it's in the include.

We go here include only object types in

the result. And so this should only have

our one categorical column which is

ocean proximity. So, housing cat is

going to have a reference to our uh it's

going to be a list that has a a

basically just our ocean proximity

feature because this select dtypes will

make sure we only pick object types and

um

grab those columns. So this is a way to

neatly grab um our categorical features

here by including the object types. Now

on the flip side we can exclude object

types and get everything else. So this

is going to be all other columns which

is excluding the object. So this is

excluding this meaning we should get all

of our numerical features that way. So

this will be all of our numericals

by excluding the object type and this

will be our housing num which is short

for numerical. So this excludes

the uh object type meaning all numerical

features

right all numerical features there.

Okay.

So, if we were to uh let's double check

this. Let's sanity check this. If we

were to print out the housing

cat, um this should be just the ocean

proximity feature, which it is. So, just

that one. If we were to print out the

housing num, this should be all the

numerical features, which are all these

guys. So it's just a reference to those

columns so that we can uh use those

later when we're mapping uh this

transform needs to go to this column

like the one hot encoding needs to go to

this column and the scaling needs to go

to these columns right so we have those

uh names of those columns already at our

disposal. So, we're just doing that.

And this is just a

simple check.

Uh, are you guys able to run this?

If you're following along, let me pause

there. Make sure I'm not going too fast.

Uh it so the the issue with a specific

data type like that is none of these are

ants. They're actually all floats. So we

did float. I think that should work. But

yes, that's the idea.

Great. I'm glad to hear that right there

with me. Great. Glad to hear that.

Okay. So, we have our columns picked out

here, which we're going to use later.

Okay.

All right. So, let's go ahead and build

out our steps for each of these types.

So, um for our numerical features, let's

build out our pipeline steps. So what

we're going to do is build out a

numerical pipeline. And it's going to be

a pipeline with a list

of tupils. And the reason these are

tupils is because every tupil has a

name. So here this is a name that we can

it can be whatever we want it to be. So

we're calling it imputer. We could call

it anything we want. We could call it

fill in the blanks. We could call it

null filling. Call it whatever you want.

We're calling it imputer because that's

that's a pretty um easy name for it. An

accurate name to what it's doing. Um but

the important thing is after the name

you give it, you put in the scikitlearn

object that you are going to use to

operate on your data. So in this case,

we're using a simple impery

of median. Now that's a choice. We could

use a strategy of mean, max. Um, we

could provide it a constant default

value. But what this means is we are

going to fill any blanks we find in

those columns with the median value of

that column. That's the strategy for the

computer. So that's pretty cool. This is

kind of an automated way to fill in the

blanks using for any column using its

median,

right? And so we could change that. We

could put mean here or max or min or

whatever. Um

but we are filling in the blank on any

column with its median. And the reason

this works is because we are going to

apply this pipeline only to these

numerical features. So that is fine.

We're we're not going to apply it to the

categorical features. We're going to

apply it to only those numerical. So it

should have a median value, right? So

that that's totally fine. So we're going

to now look at how we're constructing

the steps. We have a list of tupils.

Here's one tupole

which is the imputer with a simple imper

of strategy median. And then we can have

as many tupils as we want which

represent processing steps. So every let

me write that down. Every tupil

represents

a pre-processing

step on our data.

Okay, so we have an imputer step named

imputer and the reason it has a name is

just so you can reference it in the

pipeline if you need to. So you so it

has like a a reference name um that you

give it. Um but this is the more

important part is the actual scikitlearn

object that's doing the processing. So

in this case a simple computer but

notice that we have a secondary step

which is our scaling. Now this makes

sense. This is something we should be

doing to our features is we should be

scaling them. So here we we say okay

let's fill in any blanks first.

By the way order

matters.

So, and what I mean by that is the

simple imputer

is before the scaler. Now, that's

important because what that means is we

should be filling in any blanks before

we attempt scaling.

So, that order actually matters. We're

going to fill in blanks first in this

list. That's first. We're going to fill

in blanks. Then we are going to scale

right then we scale which makes sense

right so we we fill in blanks first then

we apply the scaler to scale our

features so those are our two steps

so so pretty simple um we are building

out our two steps now this is just one

piece of the puzzle we are going to put

this pipeline together with our one hot

encoding that's going to be coming up

next and build out our final pipeline.

But this is um a a pipeline that has two

steps that will actually be used with a

larger pipeline coming up where we we do

one hot encoding to our categoricals and

then we put a model in there at the end

to train and and use for prediction. So

um pipelines can actually be composed is

is uh something to realize there is that

we can have a pipeline that contains a

few steps. We can have another pipeline

over here that contains a few steps and

we can actually um kind of put them

together into a final pipeline that has

both pipelines uh kind of merged

together. Okay. So we're going to see

that coming up when we construct our

final one. Our final one, as you can

imagine, needs to handle this mapping of

basically saying, let's do one hot

encoding to these guys and then do this

pipeline here to these numerical

features. That's what our final pipeline

needs to handle. And it will. We're

going to build that out.

But let me pause here. Um, were you guys

able to run this? Are you with me on

this this pipeline here?

Does that make sense? Those two steps

one is filling in blanks with a median

whatever column. So where so this is

this is what's so amazing about this is

this is going to automatically search

for nulls and if you come across a

column with a null, it's going to use

the median of that column

to fill in the blank, right? To fill in

those nles.

Okay,

great. Glad to hear. Glad to hear.

Okay.

All right. So we are going to now um put

this together with a column transformer

to basically say what steps are going to

be mapped to what columns.

Um so now you can see what we're doing

here is using the column transformer

which is going to be a list of tupils

again. So this is another um list of

tupils.

But the important thing is um

each tupil

has a name

followed by so it has a name uh which

again is is generic. You can say

whatever you want it to be. So here

we're kind of shortening this to

numerical. This is short for

categorical. But the important thing is

it's followed by a pipeline

slashstep

followed by a pipeline slashstep

um followed by a uh followed by a list

of columns that it applies to. So you

can see that pattern here. What we're

saying is we're going to apply that

numerical pipeline we just defined. So

this is saved in a numerical pipeline

object here. We're going to apply that

to those numerical features. So this is

that list

of numerical features here. So that's

how we do the mapping. We have a tupil

here that says okay apply these steps to

these columns.

Those go together in that tupil, right?

Apply these steps to this uh these

columns. And then apply this step. Now

what is the step? This is a one hot

encoder

which is going to uh uh encode um those

features and it's going to uh ignore um

basically nulls for now. That's a choice

but it's going to ignore um uh basically

ignore nles and and uh skip over them

for now. We now we know there's no NLES

because we already did an is NA from

before and we know there's not any NLES

in that ocean proximity. So this isn't

going to be an issue. But that's what

that would do.

But we have a one hot encoder here which

we're going to apply to our categorical

features. Now of course that's just the

ocean proximity feature but that but

again you see the pattern in the tupole

is apply this transform which is the one

hot encoding to this column apply these

numerical transforms which is a whole

pipeline. So it's two steps in a

pipeline of um

uh an imputer and a scaler are going to

be applied to this

really nice. So those are going to be

all together in this column transformer

and that is our way to signal that for

these numerical features use these

steps. For our categorical features use

this step and and you know if we had

more than one step we were applying to

categorical we could build a pipeline

for the categorical and it would and do

the same thing. We have more than one

step here and so it's good practice when

you have more than one step to just put

that in a pipeline because we have more

than one step. We'll just put that in

this list inside of the pipeline and we

can map that pipeline to those features.

Here we only have one step. So it's okay

to just put that there um and apply that

to the categorical features. But if we

had more than one step um it would be

good practice to put that in a pipeline

which is what we do here. Right? This

pipeline is being mapped to these

features. This step is being applied to

this feature.

Okay,

how about that? Are you guys able to run

that one? Does that make sense what we

have set up so far? So, we're almost

there. We almost have our final

pipeline. We have our pre-processing

basically done to say our numerical

features should be processed with that

other pipeline and our categorical

features should be one hot encoded.

We're getting close. The only thing

we're really missing here is a model.

The only thing we're really missing is

to have our final model training

pipeline is to actually include a model

which should come at the end.

Right? So it should we should be doing

these steps first

then doing modeling which we know right

we we've done that uh many times. We've

done our pre-processing and then we do

our modeling.

Any questions on that?

Okay.

Fantastic.

All right.

So, if we wanted to uh see if we wanted

to test this so far, um we could. So we

could run the pre-processing and

actually run a fit transform on our data

and this will um basically apply that

pipeline to the data. Now this would be

a sanity check. This is a good this is a

good kind of um this is a good sanity

check that our pre-processing

works. So it's doing what we expected to

do. It's not our final pipeline because

we don't have our model in there yet.

But this is just to ensure that all of

the features are kind of behaving as we

expect. So we can uh we can do that and

we can take a look at the um results.

This looks pretty good. This all of our

numerical features ended up scaled

which is pretty good. And we have one

hot encoded features for that ocean

proximity over here.

Okay. So this looks pretty this looks

reasonable of those steps being applied

to the right columns. But this is a good

kind of sanity check to just run our fit

transform on our data to ensure those

steps are actually happening and they

are. You can see here the result of the

scaling and the uh the one hot encoding.

So that that all looks pretty

reasonable,

right?

And uh what we should also do is make

sure there are no nulls in this which

there shouldn't be because we did the

imper. So we should be doing uh is na

dot

sum

and there is no nulls anymore. So that

looks pretty good right? Those got

filled in uh by doing our steps. our

pipeline steps executed really nicely on

our training data um and and we were off

and running. And there's nothing unique

about the training data. We could do

this to our test data as well

and verify that those steps are running

and they would, right? There's nothing

really that special about running it on

the training data. Um it should also

work on the test features as well and it

does. You can check that for yourself.

Okay.

All right. So, that's pretty cool. We

can uh verify all that's working.

Any questions on that?

We're almost there with our full

pipeline. This this is this is not the

full pipeline, but this is something

that will run during our full pipeline.

Of course, our features are going to be

transformed according to those steps and

then it will be uh put into our model to

either predict or train with. Um

so let's do that. Let's actually build

out our final uh model here. So it's

actually going to be really easy to do.

All we need to do is um put in our

model. So here we're going to import the

ridge model here. Now, we could use any

we could use linear regression, we could

use lasso, we could use elastic net. Um,

we're just going to use ridge um uh um

just to test it out. And um we are going

to uh now put in a final pipeline. So,

we're going to use our pipeline. And so,

we're going to create a new one here and

map our pre-processing to our

pre-processing that we've already built.

So this is a column transformer that

already has all of our steps. And then

notice what comes after it is just the

model. Now that's pretty pretty basic,

but it makes sense that it should come

after that model. Um and of course this

is a generic name. We could we can name

it whatever we want to. Um model ridge

is pretty reasonable um to because it is

a ridge uh regression. But uh of course

we could we could change that.

Okay. So that builds out our uh final um

pipeline. So now we have a pipeline. And

what's great about that is this signals

that all of these steps should be

completed prior to doing anything with

this model. So all of those processing

steps are going to run and then we're

going to do ffit or predict and that. So

that's really great. It ensures that all

those steps are running together every

single time we call.predict with this

with this model. So we're just going to

use the pipeline in place of the model

to ensure that all those steps are

running together. And this is our this

is kind of our final pipeline that we

would use uh with like something like

ffit or predict.

So let me make that a note of that. Now

we can use this final pipeline just like

a regular model i.e. pipeline.fit

or pipeline.predict.

So we could use it in ei in either

fashion uh to to train the pipeline

would be this guy and then use the

pipeline to predict would be this. And

what we should realize is under the

hood, these steps are running first and

then we train it or these steps run

first then we use it for prediction.

Okay,

questions on that. Does that make sense

on this final pipeline here? It's just

now it it's really cool because we have

a pipeline

made up of a of a pipeline really,

right? a pipeline made up of a pipeline.

But that's scikitlearn allows you to do

that to compose pipelines in this way.

That's is pretty uh pretty uh normal

there.

Okay.

What I want to show you is we can

actually use this pipeline in a grid

search. So that's pretty amazing. We can

use this pipeline in any way we can use

a mo like a regular model. It's just

that now our pre-processing steps have

kind of been packaged together with our

model to ensure that they always run

anytime we do any processing with this

model. Um so for instance we can do a

grid search just like we did with a

regular with with just a model right

with just this. Um we can do the same

thing with the whole pipeline. Um, so

the only catch is that you want to make

sure in your grid you name things in the

appropriate way inside of your your uh

keys in your dictionary. So uh for

instance um inside of the grid uh we're

going to set up the alpha that would be

used with this ridge regression by

referencing its name. So this is model

ridge is this is the name of the model

inside of the pipeline. So you want to

make sure that goes first.

And then what scikitlearn does is it

recognizes parameters that belong with

this model by using a double underscore.

So the so you have underscore alpha um

here. So the double

uh underscore

signals a parameter

belonging to model ridge in the in the

pipeline.

Okay, so we have a model ridge is just a

reference to the model in our pipeline.

That's the one we're going to test out

these parameters with. and

underscore_pha is just a way to say this

alpha belongs to this model. Okay, it

belongs so it's going to be used with

that model in our pipeline. Um otherwise

it's going to work exactly the same way.

It's just we need to line up this naming

convention of of scikitlearn.

You just have to reference this to

whatever name you provided here and then

underscore parameter. So L1 ratio alpha

whatever right would go there.

Okay. So there is a range from 0.1 to2

uh step size of 0.1

um and then we do our grid search CV. So

this is exactly the same setup as we had

before. It's just that our model is now

the pipeline. So our pipeline is going

in there. Um we have our grid going in

there. we have our scoring is the same,

you know, negative absolute error. Um,

we're using five-fold cross validation

and we're parallelizing that search. Um,

so we're going to search through these

alphas and uh basically fit this to our

um data and find the best um find the

best alpha.

So, it's going to try out all those

combinations and try to come up with the

best alpha.

So looks like the best alpha was 0.1 for

the ridge.

Okay. Is the best. So then um if we

wanted to we could uh then predict using

the model um which would be doing

something like this. Um, and we could

also go back and do something like so we

could

now use um this param. So we could do

model

um equals ridge

and then we could put in our alpha.

Um, alpha is our results, our best

parameters, and then we get that model

ridge alpha. And then we just rebuild

our our pipeline

equals um pipeline and then we uh put in

this new model here. So we could do

this. This would be going back and just

um putting in our best alpha here for

this model and then uh ensuring that's

part of our our pipeline. So we're just

overwriting that pipeline with the best

model there

to get the best model in our pipeline.

Okay,

so that's all this is doing is just

initializing a new um let me actually I

can put this code in here.

This is actually just getting this is

just getting a model with the best alpha

and then reinserting that into our our

uh we're just overwriting our final

pipeline there with the best model that

we have.

So pretty cool that pipeline can be used

basically exactly like a model, right?

It's it's going right here in the grid

search and being used uh entirely like a

basic model. So we do ffit

um and that allows us to use it. We

could do predict we could even do

pipeline.predict once we we could go

back and do final pipeline.fit

um with this and then final

pipeline.predict with this and evaluate

Okay,

so pretty cool that pipeline can be used

uh basically exactly like how a model

would be any way we'd use a model.f

model.pred predict we can use a

pipeline.

So grid search is for instance something

that can use a model in there. Um but

instead of just a model we're ensuring

we have our pre-processing steps kind of

bundled with that model in this

pipeline.

Any

questions on

uh this example so far?

Were you guys able to run it up to here?

Were you able to run the grid search?

Okay, great.

Okay.

Okay. So, this is this is uh just

showing you what's actually happening

underneath the hood is uh you know,

we're doing some scaling. We're doing

some one hot encoding. Um

and we're doing some uh we're doing a

model here. And that's all part of our

pipeline. Um, and then we can use the

pipeline however we want. So for

example, I know it's not here, but for

an example, we could use um once we do

once we have this final pipeline um we

can can use the final um

pipeline to predict. So we can do um

predictions

equals final

pipeline.predict

and then we can pass in our test data.

Now what happens on this is once we have

ran our our pipeline.fit we have a

trained pipeline and then when we run

this final pipeline.predict uh this data

is going to be transformed.

It's going to go through those

transformation steps and then we would

apply our model to it at the end uh to

to make those predictions and then we

can evaluate those predictions which is

what we're doing kind of here

right.

Okay.

All right. So in conclusion uh we have

gone through a lot of stuff here. Um,

we've gone through regression, we've

done the regularization on regression.

So hopefully we have a good foundation

on regression. Um, what we're going to

do in a little bit is actually do some

additional practice with regression on a

new problem. We're going to do a

capstone problem and do some additional

regression work with that. Um, so we'll

do that next. Um but the other thing we

learned is how to evaluate the

regression using things like mean

squared error, RMSSE which is square

root of that. Um which is which is

really cool. So we have a sense of that

error which is our distance from our

prediction to the actual value. That's

always what these uh that's always what

these things are doing like this, right?

This mean absolute error metric from

scikitlearn is computing the average

distance from these predictions to these

test labels that we have right those

actual values. Um and that gives us a

sense on average how far away are our

predictions

um to see how good of a model that we

have, right? And we should be evaluating

that error generally

um against the scale of our targets to

see, you know,

uh how far off we typically are.

Okay. Any questions at all on this

lesson on regression? Uh anything we

covered up to this point?

We're going to do some more practice

with the next. We'll do the capstone. So

we get So we just do some more

regression problems.

Yeah, it's a that's another bad score.

It's a little bit hard to interpret this

though because it's m ae. Um so one

thing we could do is is compute mean

squared error and then take the square

root of it to get the RMSSE which is a

much better uh evaluation metric in

terms of our target. Um so we could

actually run that. Uh if we go back here

and um we could generate for instance we

could generate the MSE which is the mean

squared

error

and it's it's the same exact function uh

of using our predictions

um

and then we could just print that out

mean squared error.

So we have mean squared error and then

what we can do is let's take the um MP.

Square root of that.

So that way we can generate the RMSSE.

So yeah that I mean that's pretty bad.

That's uh pretty bad. Uh now let's let's

go back and look at our

uh data though. So let's take a look at

the average for our Y. Um remember one

thing we should be doing is taking a

look at um what our uh let's take a look

at y test mean

to get an average value. So the average

value is in the 200,000s. So

this isn't this isn't awful. This is

70,000. It's still a decent amount of

error. It's not as bad as the models we

have before though, right? This is an

average

median price of the house is in the

26,000 range and our error is off by

like 70,000,

right?

So, it's not good. Um, but it's not

hor like as bad as the it's not as

horrible as we've seen so far. Right.

This is a little bit better of a model.

A little bit better. closer to zero

would be better, right? Um but the

smaller the better. But uh remember this

is the um these even the mean absolute

error is is technically in similar units

as the as the uh um

as the target. So 50,000 60,000 here

70,000 it's still a decent amount of

error in terms of 200,000.

Uh so far we only come up with models

and test their accuracy with available

data. We haven't used a model to make

completely new predictions on No, we

haven't done that. Uh except we know how

to do that. Um it would so to make

predictions on new data would be exactly

how we're making them on our available

data because we actually do that all the

time. If we go back down to our model

building,

um it's it looks just like this, right?

where we take so for instance we do

predictions all the time on test data

that was never involved in the training.

So it's it's as if this data mimics new

data that we've never seen before. So if

we had new raw data it would just it

would be the same exact process. the new

now with our pipeline it makes it a

little bit easier because with the

pipeline

um the raw data will go through those

transformations which it should right

the raw data should because if it's

missing data it needs to be filled in if

it has categoricals it needs to be one

hot encoded so that's the purpose of the

pipeline actually is to make sure that

if we're dealing with raw data um those

steps can happen on the data before it

goes into the model. Right? So we so

that's kind of the purpose of the

pipeline

is to ensure that we run those steps

ahead of using it using a model with it.

But but ultimately that's how it uh any

scikitlearn model is going to be doing

the predict even if it's a pipeline

right it's going to be uh we just go

back down here it's going to be um

predict it's always going to be that on

new data

yeah

okay

really good question uh what I wanted to

do next was do some practice um I wanted

to go over to the capstone session five.

So, in this course, we have some more

capstone sessions. So, if you have a

moment, you want to pull up those

capstone session materials, the

incremental capstone session materials.

Um, we're going to be doing session five

today. So, this is just um remember it's

just extra practice that we do after

we've covered some concepts. So we are

going to do some regression practice now

that we've uh covered regression um

pretty fully and then um this will this

will be good practice before we head

into lesson four on classification. So I

just want to do this practice now while

it's fresh while the material is kind of

fresh in our in our minds. Um do you

guys have the capstone materials? Do you

know where to get it in your LMS? It's

in your LMS and the resources the uh

capstone materials you want to download

that so you can get the the data um and

the the slides for the instructions

right the or PDF I think for you guys

but uh let me ask you do you have those

we want to pull up session five if you

have it

thank you I was just going to share that

appreciate that yeah so this is going to

be session Question five. Um, you're

also going to want the data that we're

going to use with this, which is going

to be the, uh, bike rental data set.

I'll share that with you guys now.

So, we're going to be using this bike

rentals data set for this uh, for this

capstone. Um, it's the one we're going

to build a regression model off of.

Okay. So, you should have that one from

the the capstones data sets as well.

All right. So, let's go through this.

Um, we are going to be doing uh machine

learning here. So, we're going to be

doing uh so we're talking about that

kind of example product that Aura

product that was in our original

capstone. Um, and in order to do uh to

to to um make decisions, it's going to

have to build some models. Um, in this

case, it's going to be doing some bike

rental modeling. um which is the data

set we have. So we're moving away from

that healthcare data set going into this

bike rental data set as an example of

the capabilities here. Um so we're going

to do this first capstone. Um after we

do classification, we'll do this

practice. Um after we do unsupervised

learning, we'll do this practice on

clustering. And then after we do

recommendation uh which is the last

lesson um we'll come back and do

practice with building a recommendation

engine. Okay, but we're going to do this

one today. Um and then we will uh do

these other capstones as we go along. So

this will be session six, session seven,

and session 8

uh later on in the course. Okay.

Um,

okay. A little bit about the data. So,

uh, in this capstone, we're going to be

working with a, um, a shop, like a

retail shop that rents out bikes. And

they have data, um, on a per day basis

with, um, actually on a per hour basis

on the number of bikes that they rented

in every hour. Um, so maybe one hour

they rented out 20 bikes, another hour

they rented out 30 bikes. Um, another

hour they rented out 15. So they have

that data here in in the CSV. Um, they

have other kinds of data like the

environmental data like what the

temperature was at that hour, the

humidity, if there was snowfall if it's

a holiday, the wind, the visibility, the

due point, um, solar radiation,

rainfall, uh, what season it was. um

what uh what day it was like

um the functional is like if it's uh I

believe it's like if it's a weekend or

weekday um which would be uh

nonfunctional

um so

based on those features we have a bunch

of tasks okay so we have um based on the

uh rented by count hour of the day um

temperature, humidity, wind speed,

rainfall, and whatever other features

we're that are in the data set. We're

actually going to build a model to

predict the bike count required for

every hour to have a stable supply of

rented bikes. So, our goal is actually

going to be to predict the bike rental

count per hour. Um and uh we are going

to do that using all the features we

have at our disposal. um like mostly

those environmental features and what

day it is, those kind of things. Um

so we're going to load our data. We're

going to do our usual check. So check

for any NLES, handle those missing nles.

Um we're actually going to practice

converting our date because we actually

do have things based on a date here. So

we can convert it over to a datetime

object, extract different uh features

from that um like the month or the day

of the week. Um we're going to check uh

correlation using the heat map. We're

going to do some plots, some very basic

plots. The focus is going to be on the

modeling. So, I probably won't spend too

much time on the plots today, but um

there's some plot tasks in here like the

uh the the histogram of the bike count,

the histogram of the numerical features.

Um box plot of the bikes against the

categoricals.

Um so we can do some plots like that

from Seabor for instance. Uh Seabor

category plot of rented bike count

against features like hour, holiday,

rainfall, snowfall. Um, so we can see

how that stacks up against like

different hours of the day, different

holidays, rainfall, different weather

events. Um, then we're going to do then

we're going to start building our model,

right? So encode our categorical

features. Um, identify target variable

and do the split and then do scaling and

do three different models. So we're

actually going to build a linear

regression. We're going to build a lasso

regression and build a ridge regression

for the hourly bike count. and we're

going to see which model performs the

best. Now,

we could and should um build this into a

pipeline. So, that could be something we

practice. Um but this initial

instructions actually doesn't require

doing that, but I think it's really good

practice to um build our model. So, we

could use git dummies. Like it says

here, hint to use git dummies. We could

do that and build our model that way,

but more practical would be doing the

steps we did towards the end of lesson

three, which is um actually putting

everything together into a pipeline. All

right, that'd be more practical and then

fitting the pipeline and predicting with

it um for evaluation.

So, uh we'll do that. I think we'll do

that instead because I think that'll be

more practical. The the pipelines are

really uh useful. So the things we'll

have to do when we build our pipeline

will be um making sure we handle the

missing values. So we'll want that imper

in there for numerical features. We'll

want our one hot encoding very similar

pipeline to the one we built earlier. Um

and then we'll want to basically map

those to the right columns using the

column transformer. And then um we'll

have our pipeline ready to go for

training and prediction.

Right.

Okay. So, those are going to be our uh

steps. Any questions on this before we

kind of get started on it.

Okay. So, let me go over to the

notebooks and let me actually do a new

notebook.

And I'm going to name this uh

capstone

session five.

Okay.

So, I'm going to come back here and

reference the steps here. Okay. So the

steps are to load our data set and

basically check for nles.

So we should be pretty adept at doing

that. We're going to import pandas as

pd.

And I need to make sure the data sets

available. So I need to

uh load that here. So we're going to use

our bike rental.

Florida bike rentals.

And then look at the first five rows.

Uh, I got a decoder error.

UTF8 codec can decode by in

one moment.

I think we have an error in the data

set. Was anybody able to get this to

run?

Hopefully they

or did you get the same error as me?

giving me the same error.

I think we need to set a

encoding

having other issues. What other issues?

Sorry, I think this data is kind of

corrupted.

I may need to open it externally.

Okay, let me open it. There may be just

a bad character that needs to be

removed.

Okay,

let me try a different Let me try

something real quick.

I think the data is needs to be updated.

Oh, saved it as the wrong file.

One moment.

Okay, let me try uploading this.

Okay, there. That worked better. So, let

me give you the data. I think it was uh

yeah, I think it was that the degree

code was giving it some issues. So, I

actually just removed it.

You can do that or you could just work

with this one.

I removed the temperature had a strange

like degree symbol that wasn't being

parsed.

So, I just removed that in this data

set. This one should work. The one I

just sent you guys should work. Or yeah,

I guess you could try the CP1252

with the original data. See if that

works. Did that work for you?

Okay, it works with that encoding. So

yeah, you could use that encoding or

uh

remove that temperature degree which is

what I did from that. So I use that

other data set.

Yeah, that opens the data but

it's not it's not in a data frame.

We just we want it in a dataf frame to

work with our models and doing all of

our Yeah. Like that opens the file. If

it just if it was a text file that's

fine, but it's not in a data frame. Want

in a structured data frame so we could

use it.

Okay. So, one of those methodologies

hopefully works. So you can either work

with the file I sent and read it like

this or you can use the encoding.

Did you guys were other people able to

open it once they change the encoding?

Okay.

Okay, very good.

Let's see what we have. Let's do

df.info.

Let's see what we have, which is pretty

standard step to do once we first load

in some data. Um, so we have about 14

columns here. We have a date and then we

have um which is an object right now but

we're actually going to convert that

over to a datetime object in a minute.

Um we have our bike count which is what

we want to ultimately predict. This is

the target variable in our data.

Remember that's going to be the problem

is to predict that hourly bike count. Um

the hour of the day that we are

producing that bike count um is a

feature. the temperature which is in

degrees Celsius. Uh I I removed that

from there but that's it was in Celsius.

Um

and then we have a bunch of different

features which are uh a kind of boolean

like yes no holiday no holiday the

season. So these guys are going to be

good um candidates for

uh these are going to be good candidates

for um doing our uh one hot encoding.

Right? These are probably the three that

we should pick to do uh some

transformations to for our one encoding.

Okay.

Everything else is pretty numerical

though, so those should be fine to keep

those. But they would they're just going

to be good candidates for scaling,

right? Really good candidates for

scaling.

All right, let's go back to here.

And let's go back to the task.

So we loaded the data set. Um we're we

are going to check for nulls. It looks

like there's actually not any nles. So

there may not be any uh imputing that we

really need to do for this. Um but we

can check. So we can use is na or is

null and then do the sum.

Um looks like we don't have any. So that

that's good. There's no nulls.

That's pretty good. So we don't have to

worry about filling in any blanks

really.

Um but you know that would normally um

that would be an important part of our

pipeline right is filling in any nles.

Looks like we don't have to worry about

that here.

So that's good.

So we can

uh

we can extract we can do the date uh

extraction. And now one thing that we're

going to do is uh create purposely for

this we're going to create a weekend or

weekday feature. Okay, weekday or

weekend which should be really easy to

do from the day of the week. um which we

should be able to extract from this uh

from the date.

Let's go back to Were you by the way,

were you guys able to run this?

Just want to make sure everyone's with

me. Checking if there's any nulls. We

did that.

Okay.

All right. So, we're following along

there. We checked if there's any nles.

Let's do um our conversion. So, let's do

um df

uh date.

And this is going to be um

PD.2

date time and then df

date.

All right. And then let's sanity check

that that it got converted by doing

info.

Oh, we might need this format.

Um, maybe we should do

Okay, so that actually worked. Let's

see. Let's double check the

date.

Okay, so that worked. It extracted it

into

uh it extracted it into the right dates.

Some weird encoding with this

So it goes all the way up to 2018

from

the very beginning data is in 2017 the

beginning of the year right. So it go

and then the tail is all the way in

2018.

This this extracts the date.

If you guys want to run that

you guys able to run this

to extract the date. Okay, perfect. So,

this extracts the date. By the way, the

reason we need this is because our date

formats are not uniform. Uh they were

actually in different encodings. So some

of them had the year uh the string was

in a slightly different format where the

year was last or the year was first. So

if we do format mixed, it kind of

rearrang it kind of puts it in a uniform

arrangement with the year. It's year,

month, date, but it parses that out um

correctly. Uh so we have year, month,

date in there.

Um but the this strings were in a mixed

format. So we uh put that argument in

there to handle that case.

Okay. So the reason we're going to do

that is that we should be able to create

a day of the week uh feature. Um so

let's actually do that and add it to our

data frame. So we should be able to

create um day of week

And we should be able to extract um

the DF

date. And then we do our usual DT dot um

day

of week.

Okay. So, and then let's see what that

does. So, if we add that, let's actually

see um let's see what our new features

are.

So we add that it should go onto the end

of the data frame as the day of the week

is a numerical day of the week. So this

this is the uh um

this looks like a th or no a Wednesday.

I think that's the third day.

Sunday Monday being uh Monday being

actually this would be a Thursday. I

think Monday would be zero.

This would be Thursday.

Okay. So, we extract that day of the

week.

So, it's just we're adding a new column

called day of week, which extracts the

day of the week from the date feature,

the datetime feature, uh the day of the

week.

And we're verifying that here. It's now

a new feature called day of the week.

which is a number.

Does that make sense what this is doing?

Yeah, it's a Thursday. So, I think yeah,

Monday is zero

and Sunday is six. So,

um, Monday is

zero, Sunday is

six.

Yeah. So, this should be a Thursday.

This first five rows is a Thursday. And,

and by the way, the data um, this is all

on the Thursday. This is a different

hours of the day.

Different hours of the day. Um, and we

the the thing that we don't know about

this is the necessarily the time zone.

So, it may seem strange that like this

is zero, which is kind of like midnight.

Um, it could be it could be in a

different time zone. So, we don't know

that uh necessarily, but the there is um

you know 250 bikes, 200 bikes, 173. It

starts to decrease over these hours.

Just kind of notice that.

Okay. So, by the way, uh we should be

able to create a new feature. So let's

create

the weekend feature

which is um we can create using

uh weekend

and we can take our DF

um day of week

and then we can just um

say is this um greater than or equal to

uh greater than or equal to five

cuz that would be five or six. And we're

going to

um put this as

actually. Let's just leave it like that.

That should be fine.

No, let's change it as type

uh int.

So, let's see this feature.

Let's see if this works.

So this is not a weekend uh because it's

a Thursday, right? So anything bigger

than five would be bigger than or equal

to five would be five or six which would

be Saturday, Sunday. Um so we know it's

not a weekend day. This is uh zero.

Okay.

So just creating those features there

and I can paste these in. Does that make

sense what what I just did?

Any questions on that code? And this is

important. If you don't have this um it

this will be a true or a false, but we

want it to be a zero or a one. So when

it's actually true, this should be a

one. When it's false, it'll be a zero.

So, we want it to be an integer rather

than a true false. So, that's why I have

this part. That's why I did this here

to make sure it's um make sure it's an

integer.

Good.

Okay,

so we're pretty much uh doing that and

we convert it to day and extract day. Um

let's

check our heat map.

Let's do that.

So let's import Seabour

as SNS.

So, we're going to do our heat map next.

Let me uh make a note of that. So, we're

going to do

heat mapap

heat map. Um,

so we're going to do uh SNS

heat map

and then let's do uh df.correlation.

And let's make sure we do numeric

um only

equals to true

and then let's do annotate

equals to true

so that we get those uh correlation

values that are displayed on the heat

map. So what this is going to do is

create our heat map with our correlation

matrix um where we're making sure we

only do the numerical features of course

when we do uh the heat map.

Let's generate that. Um okay so we have

some

uh pretty mild um correlations. Now the

ones that are so the ones that are

correlated are the temperature and

dupoint temperature. Those are pretty

correlated.

Um so I think we could argue that we

should drop one of those. Probably just

the dupoint temperature we could drop.

Um that's a really high correlation

right right here.

Let me draw it in red. That's a this

this one here

which is the Dupoint temperature against

the regular temperature. That's super

high. 0.91

is nearly a onetoone correlation.

So, that's a good candidate to be

dropped. Uh, one of those guys, I would

argue probably the Dupoint temperature

we could get rid of dropping. Um, and

just keep the regular temperature

because they're nearly identical. Um, if

you go down here,

day of the week and weekend are

correlated. Um, and that makes sense.

That's a pretty strong correlation

because of course if it's depending on

what day of the week it is, it is the

weekend or not. So we could go and we

could go ahead and drop the day of the

week column if we wanted to because

we've already derived the weekend uh

feature which is a simpler feature. Is

it the weekend or is it not the weekend?

Um so we could probably drop day of week

and be okay. That's a pretty strong

correlation. Otherwise, it's all pretty

weak. I don't see any other strong

correlations

uh necessarily.

Um so

that seems pretty reasonable is that we

could get rid of we could get rid of

this one and we could get rid of day of

week and probably be okay,

right? Those are pretty strong

correlations. is 08 and 0.91.

Pretty strong.

Okay. Were you guys able to run that one

and see the the uh heat map? And does

that make sense based on what I'm

saying?

So, we want to run the heat map

and pass in that correlation. And we

want to make sure we turn this to true

to only do the numerical features. And

then this to true to show the value

on the uh on the heat map.

able to run that one.

Great.

Okay.

Um, that's a good question. Any

recommendation on how to choose between

the two? Uh,

not really. I think

I don't think it really matters. If

they're correlated to each other, then

uh including one of them,

um, including one of them should give

you the same information as the other,

especially if they're really correlated.

So, it doesn't really matter too much.

Um

the way that I would choose is to think

about like I'll give you an example in

this day of week versus weekend. Um we

could I would argue drop day of week

because the weekend is a simpler

feature. It's only is zero or one and

that's directly derived from day of

week.

So it has less um complexity to it. it's

a little bit simpler of a feature and I

think that makes it easier to work with.

Um,

but the truth is that uh we could do

both options and try them out and

evaluate the results, right? So, well to

be to be truly thorough, what we could

do is build a model where we've dropped

this one and do the evaluation and then

go back and build a model where we've

dropped this one and do the evaluation.

Right? So that's the proper way to do it

is to actually just build both models

with each one dropped and see which

performs better.

Um otherwise I I tend to prefer to go

for the simplicity whatever one has kind

of a lower range.

But the truth is if they're if they're

really strongly correlated, it's not

going to matter too much. Uh because

they're going to give you the same

information, right? Uh because they're

so strongly correlated. Like if I in

this data, like if I know the

temperature, I pretty much know what the

Dupoint temperature is going to be.

They're so correlated.

So it doesn't it doesn't really matter

which one I drop

but I yeah prefer to go for the

simplicity.

Okay.

Um going back to let's let's do this

plot now of the distribution of the

rented bike count. So that should be

taking a look at the

um

the rented

bike count and then just doing a

histogram

and we can see what that distribution

is. Most of it is it less than 250.

Um but there are some values that are

really high, right? There are some days

that turn out to be over 3,000 into the

3500 range. Um, that's quite a bit, but

most of the days are stacked over here

in this like 250 bucket. So, by far

that's the most. And then it kind of

decreases from there. Most values are in

that range and then uh it kind of

declines.

So that should just be this simple

histogram here.

So this is the

histogram.

That should be a simple one to build.

Were

you guys able to run that one?

Sweet.

Okay, pretty basic. Just showing us how

it's distributed. And we kind of noticed

that most of it is in the 250 bucket or

below. But there's a good amount that's

out there beyond like in the 500,

7500,000

um

all the way up to there's some days that

register with a 3,000 and above, right?

3500.

So,

in fact, we could look, we didn't do

this, but it might be worth doing is we

could look at the describe because we

never looked at the maximum.

Um,

so for the bike count, there are some

days that are zero

and the maximum is 3500. 3556

and the median is around 500 bikes.

Um

the average around 700.

So that's just our usual describe

All right, let's see what else. Uh,

plot the histogram of all numerical

features. So, uh, let's do that.

So uh luckily we have a shortcut to do

this. If you guys remember we have our

SNS um pair plot and we can pass in our

uh our data

is our df right so we can pass that in.

Um so this is kind of a nice this is a

nice thing to run that will um generate

the uh the scatter plots of all the

features against each other kind of like

the correlation but at the same time

produce the histograms on the diagonal.

Right? So this should be a nice plot to

to see all the histograms

uh in one kind of uh grid.

That's just this one.

Okay, so pretty this is obviously quite

a bit of data, but this is all the

features against each other. Um, look at

this. I mean, a couple things you see

right away is look at the ones that are

really highly correlated like the

temperature against the Dupoint

temperature. Do you guys see this strong

correlation here?

So that's that's a very indicative of a

very strong correlation, right? That's

the temperature against the Dupoint

temperature. That kind of makes sense

that it's uh that was the 0.91

correlation. So of course it's like

that.

Of course it looks like that, right? Uh

just a very strong correlation on the on

the x-axis or sorry on the diagonal is

all the histograms. They're a little bit

zoomed out so difficult to see. Um but

um we can get a sense of how some of

these are distributed like the first one

uh

sorry this first one

which is the bite count we already did

um temperature we can see how that's

distributed this third one it's kind of

evenly distributed

left of due versus temperature this one

that one is the hours

versus the due point. So, it it's kind

of evenly spaced out. Um, which would

probably be which would make sense

because this data is across two years.

So, you're going to get a lot of

seasonal data in there, right? So, it's

going to it's going to vary across like

seasons. Yeah. So, it kind of looks it's

very evenly spread out.

Okay. Were you able to get Parpot to

run?

Takes a moment to run.

takes a moment, but it produces all of

these plots, including all the

histograms.

And again, if we wanted to zoom in on

any one particular histogram, we could

do that. We just have to basically copy

and paste this code and swap out this

feature. Just swap out this feature and

we can get a zoomed in uh plot of any

one of those uh features like the

temperature

um or visibility or whatever it is.

So, what I want to do, uh, I think we'll

take one more break here. Um, and then

what we'll do is we'll come back and

just finish up our practice. Um, I'm

going to start building the model. I

know it wants us to do some additional

plotting. Um, but I want to get to

mainly the plotting elements, including

doing the pipeline one more time. Um, so

we're going to do that. Um, we're going

to practice building out the pipeline in

a similar fashion to exactly how we

built the pipeline earlier and then do

we're going to build our models and

we're going to practice that coming up.

I'll probably leave the plotting to you

guys to do as kind of homework if you

want to do that. Um, just because I want

to get to the to the modeling. Hello and

welcome to machine learning tutorial

part one. This is part one of a machine

learning series put on by SimplyLearn.

My name is Richard Kersner. I'm with the

SimplyLearn team. That's

www.simplearn.com.

Get certified, get ahead. What's in it

for you today? Well, we'll start off

with a brief explanation of why machine

learning and what is machine learning.

And then we'll get into a few of the

types of machine learning. machine

learning algorithms, linear regression,

decision trees, support vector machine,

and finally, we'll do a use case where

we're going to classify whether a recipe

is of a cupcake or a muffin using the

SPM or the support vector machine.

Sounds like a delicious way to explore

machine learning. So, why machine

learning? Why do we even care about

having these computers come up and be

able to do all these new things for us?

Well, because machines can now drive

your car for you. still very in the

infant stage but it's just exploding as

we see with uh Google's Whimo and then

Uber had their program which

unfortunately crashed. They know that

this is huge. This is going to be the

huge industry to change our whole

transportation infrastructure. Machine

learning is now used to detect over 50

eye diseases. Do you know how amazing

that is to have a computer that doublech

checkcks for the doctor for things they

might miss? That's just huge in the

health industry. pretty soon they

actually do already have that with in

some areas where maybe not for eyes but

for other diseases where they're using

the camera on your phone to help

pre-diagnose before you go in and see

the doctor. And because the machine can

now unlock your phone with your face, I

mean, that's just cool having it being

able to identify your face or your voice

and be able to turn stuff on and off for

you depending on where you're at and

what you need. Talk about an ultimate

automation our world we live in. And as

we dig in deeper, we have a nice example

of Facebook. As you can see here, they

have the Facebook post with Halloween.

Comment yes if you want it order here.

Nobody likes spam posts on Facebook that

annoy them into interacting with likes,

shares, comments, and other actions. I

remember the original ones were all if

you don't click on here, you will have

bad luck or some kind of fear factor.

Well, this is a huge thing in a social

media when people are getting spammed.

And so this tactic known as engagement

bait takes advantage of Facebook's

newsfeed algorithm by choosing

engagement in order to get the greater

reach. To eliminate engagement bait, the

company reviewed and categorized

hundreds of thousands of posts to train

a machine learning model that detects

different types of engagement bait. So

in this case, we have we're using

Facebook, but this is of course across

all the different social media. they

have different tools are building and

the Facebook scroll gif will be replaced

kind of like a virus coming in there and

notices that there's a certain setup

with Facebook and it's able to replace

it and they have like vote baiting react

baiting share baiting they have all

these different these are kind of

general titles but there certainly are a

lot of way of baiting you to go in there

and click on something so they fed all

this this data was fed into the machine

and then they have the new post the new

post comes up that takes over part of

the Facebook setup up and that's what

you're looking at. You're looking at

this new post that's replaced like a

virus has replaced that. So what

Facebook did to eliminate this is they

start scanning for keywords and phrases

like this and checks the click-through

rate. So it starts looking for people

who are clicking through it without even

looking at it or clicking through it and

it's not something that normally would

be clicked through. Once Facebook has

scanned for these keywords and phrases,

it is now able to identify the spam

coming in and this makes your life

easier. So you're not getting spammed.

It's not like walking through an airport

and in a lot of countries you have like

hundreds of people trying to sell you

time share. Come join us. Sign up for

this. Eliminates that annoyingness. So

now you can just enjoy your Facebook and

your cat pictures. Or maybe it's your

family pictures. Mine is family.

Certainly people like their cat pictures

too. Another good example is Google's

Deep Mind project Alph Go. A computer

program that plays a board game Go has

defeated the world's number one go

player and I hope I say his name right.

Kijiji the ultimate go challenge game a

three of three was on May 27th 2017 so

that was just last year that this

happened and what makes this so

important is that you know go is just is

a game so it's not like you're driving a

car or something in our real world but

they are using games to learn how to get

the machine learning program to learn

they want it to learn how to learn and

that is a huge step a lot of this is

still in its infant stage as far as

development

as we saw what happened with the as I

referred to earlier the Uber cars. They

lost their whole division because they

jumped ahead too fast. So still an

infant stage, but boy is this like the

beginning of just an amazing world that

is automated in ways we can't even

imagine what tomorrow's going to look

like. We've looked at a lot of examples

of machine learning. So let's see if we

can give a little bit more of a concrete

definition. What is machine learning?

Machine learning is the science of

making computers learn and act like

humans by feeding data and information

without being explicitly programmed. And

we see here we have a nice little

diagram where we have our ordinary

system, your computer nowadays, you can

even run a lot of the stuff on a cell

phone because cell phones have advanced

so much. And then with artificial

intelligence and machine learning, it

now takes the data and it learns from

what happened before and then it

predicts what's going to come next. And

then really the biggest part right now

in machine learning that's going on is

it improves on that. How do we find a

new solution? So we go from descriptive

where it's learning about stuff and

understanding how it fits together to

predicting what it's going to do to

postcripting coming up with a new

solution. And when we're working on

machine learning, there's a number of

different diagrams that people have

posted for what steps to go through. A

lot of it might be very domain specific.

So if you're working on photo

identification versus language versus

medical or physics, some of these are

switched around a little bit or new

things are put in. They're very specific

to the domain. This is kind of a very

general diagram. First, you want to

define your objective. Very important to

know what it is you're wanting to

predict. Then you're going to be

collecting the data. So once you've

defined an objective, you need to

collect the data that matches. You spend

a lot of time in data science collecting

data and the next step preparing the

data. You got to make sure that your

data is clean going in. There's the old

saying, bad data in, bad answer out or

bad data out. And then once you've gone

through and we've cleaned all this stuff

coming in, then you're going to select

the algorithm. Which algorithm are you

going to use? You're going to train that

algorithm. In this case, I think we're

going to be working with SVM, the

support vector machine. Then you have to

test the model. Does this model work? Is

this a valid model for what we're doing?

And then once you've tested it, you want

to run your prediction. You want to run

your prediction or your choice or

whatever output it's going to come up

with. And then once everything is set

and you've done lots of testing, then

you want to go ahead and deploy the

model. And remember I said domain

specific. This is very general as far as

the scope of doing something. A lot of

models you get halfway through and you

realize that your data is missing

something and you have to go collect new

data because you've run a test in here

someplace along the line. You're saying,

"Hey, I'm not really getting the answers

I need." So, there's a lot of things

that are domain specific that become

part of this model. This is a very

general model, but it's a very good

model to start with. And we do have some

basic divisions of what machine learning

does that's important to know. For

instance, do you want to predict a

category? Well, if you're categorizing

thing, that's classification. For

instance, whether the stock price will

increase or decrease. So in other words,

I'm looking for a yes no answer. Is it

going up or is it going down? And in

that case, we'd actually say, is it

going up? True. If it's not going up,

it's false, meaning it's going down.

This way, it's a yes, no. 01. Do you

want to predict a quantity? That's

regression. So remember, we just did

classification. Now we're looking at

regression. These are the two major

divisions in what data is doing. For

instance, predicting the age of a person

based on the height, weight, health, and

other factors. So based on these

different factors, you might guess how

old a person is. And then there are a

lot of domain specific things like do

you want to detect an anomaly? That's

anomaly detection. This is actually very

popular right now. For instance, you

want to detect money withdrawal

anomalies. You want to know when

someone's making a withdrawal that might

not be their own account. We've actually

brought this up because this is really

big right now. If you're predicting the

stock whether to buy stock or not, you

want to be able to know if what's going

on in the stock market is an anomaly,

use a different prediction model because

something else is going on. You got to

pull out new information in there or is

this just the norm? I'm going to get my

normal return on my money invested. So

being able to detect anomalies is very

big in data science these days. Another

question that comes up which is on what

we call untrained data is do you want to

discover structure in unexplored data

and that's called clustering. For

instance, finding groups of customers

with similar behavior given a large

database of customer data containing

their demographics and past buying

records. And in this case, we might

notice that anybody who's wearing

certain set of shoes goes shopping at

certain stores or whatever it is. are

going to make certain purchases. By

having that information, it helps us to

market or group people together. So then

we can now explore that group and find

out what it is we want to market to them

if you're in the marketing world. And

that might also work in just about any

arena. You might want to group people

together whether they're uh based on

their different areas and investments

and financial background, whether you're

going to give them a loan or not. before

you even start looking at whether

they're a valid customer for the bank,

you might want to look at all these

different areas and group them together

based on unknown data. So, you're not

you don't know what the data is going to

tell you, but you want to cluster people

together that come together. Let's take

a quick detour for quiz time. Oh, my

favorite. So, we're going to have a

couple questions here under quiz time

and um we'll be posting the answers in

these part two of this tutorial. So,

let's go ahead and take a look at these

quiz times questions and hopefully

you'll get them all right and it'll get

you thinking about how to process data

and what's going on. Can you tell what's

happening in the following cases? Of

course, you're sitting there with your

cup of coffee and you have your checkbox

and your pen trying to figure out what's

your next step in your data science

analysis. So, the first one is grouping

documents into different categories

based on the topic and content of each

document. Very big these days. you know,

you have legal documents, you have uh

maybe it's a sports group documents,

maybe you're analyzing newspaper

postings, but certainly having that

automated is a huge thing in today's

world. B, identifying handwritten digits

in images correctly. So, we want to know

whether uh they're writing an A or

capital A, B, C, what are they writing

out in their hand digit, their

handwriting. C behavior of a website

indicating that the site is not working

as designed. D, predicting salary of an

individual based on his or her years of

experience with HR hiring uh setup

there. So stay tuned for part two. We'll

go ahead and answer these questions when

we get to the part two of this tutorial

or you can just simply write at the

bottom and send a note to SimplyLearn

and they'll follow up with you on it.

Back to our regular content. Now these

last few bring us into the next topic

which is another way of dividing our

types of machine learning and that is

with supervised unsupervised

and reinforcement learning. Supervised

learning is a method used to enable

machines to classify predict objects,

problems or situations based on labeled

data fed to the machine. And in here you

see we have a jumble of data with

circles, triangles and squares. And we

label them. We have what's a circle,

what's a triangle, what's a square and

we have our model training and it trains

it. So we know the answer. Very

important when you're doing supervised

learning, you already know the answer to

a lot of your information coming in. So

you have a huge group of data coming in

and then you have new data coming in. So

we've trained our model. The model now

knows the difference between a circle, a

square, a triangle. And now that we've

trained it, we can send in in this case

a square and a circle goes in and it

predicts that the top one's a square and

the next one's a circle. And you can see

that this is uh being able to predict

whether someone's going to default on a

loan because I was talking about banks

earlier. Supervised learning on stock

market whether you're going to make

money or not. That's always important.

And if you are looking to make a fortune

in the stock market, keep in mind it is

very difficult to get all the data

correct on the stock market. It is very

uh it fluctuates in ways you really hard

to predict. So it's quite a roller

coaster ride. If you're running machine

learning on the stock market, you start

realizing you really have to dig for new

data. So we have supervised learning.

And if you have supervised, we need

unsupervised learning. In unsupervised

learning, machine learning model finds

the hidden pattern in an unlabeled data.

So in this case, instead of telling it

what the circle is and what a triangle

is and what a square is, it goes in

there, looks at them, and says for

whatever reason, it groups them

together. Maybe it'll group it by the

number of corners. And it notices that a

number of them all have three corners, a

number of them all have four corners,

and a number of them all have no

corners. And it's able to filter those

through and group them together. We

talked about that earlier with looking

at a group of people who are out

shopping. We want to group them together

to find out what they have in common.

And of course, once you understand what

people have in common, maybe you have

one of them who's a customer at your

store, or you have five of them are

customer at your store, and they have a

lot in common with five others who are

not customers at your store. How do you

market to those five who aren't

customers at your store yet? They fit

the demographs of who's going to shop

there, and you'd like them to shop at

your store, not the one next door. Of

course, this is a simplified version.

You can see very easily the difference

between a triangle and a circle, which

is might not be so easy in marketing.

Reinforcement learning. Reinforcement

learning is an important type of machine

learning where an agent learns how to

behave in an environment by performing

actions and seeing the result. And we

have here where the in this case a baby.

It's actually great that they used an

infant for this slide because the

reinforcement learning is very much in

its infant stages. But it's also

probably the biggest machine learning

demand out there right now or in the

future. It's going to be coming up over

the next few years is reinforcement

learning and how to make that work for

us. And you can see here where we have

our action. In the action in this one,

it goes into the fire. Hopefully, the

baby didn't it's just a little candle,

not a giant fire pit like it looks like

here. When the baby comes out and the

new state is the baby is sad and crying

because they got burned on the fire. And

then maybe they take another action. The

baby's called the agent because it's the

one taking the actions. And in this

case, they didn't go into the fire. They

went a different direction. And now the

baby's happy and laughing and playing.

Reinforcement learning is very easy to

understand because that's how as humans

that's one of the ways we learn. We

learn whether it is you burn yourself on

the stove, don't do that anymore. Don't

touch the stove. In the big picture,

being able to have machine learning

programming or an AI be able to do this

is huge because now we're starting to

learn how to learn. That's a big jump in

the world of computer and machine

learning. And we're going to go back and

just kind of go back over supervised

versus unsupervised learning.

Understanding this is huge because this

is going to come up in any project

you're working on. We have in supervised

learning, we have labeled data. We have

direct feedback. So someone's already

gone in there and said, "Yes, that's a

triangle. No, that's not a triangle."

And then you predict an outcome. So you

have a nice prediction. This is this

this new set of data is coming in and we

know what it's going to be. And then

with unsupervised trading, it's not

labeled. So we really don't know what it

is. There's no feedback. So, we're not

telling it whether it's right or wrong.

We're not telling it whether it's a

triangle or a square. We're not telling

it to go left or right. All we do is

we're finding hidden structure in the

data, grouping the data together to find

out what connects to each other. And

then you can use these together. So,

imagine you have an image and you're not

sure what you're looking for. So, you go

in and you have the unstructured data.

Find all these things that are connected

together and then somebody looks at

those and labels them. Now you can take

that label data and program something to

predict what's in the picture. So you

can see how they go back and forth and

you can start connecting all these

different tools together to make a

bigger picture. There are many

interesting machine learning algorithms.

Let's have a look at a few of them.

Hopefully this gave you a little flavor

of what's out there and these are some

of the most important ones that are

currently being used. We'll take a look

at linear regression, decision tree and

the support vector machine. Let's start

with a closer look at linear regression.

Linear regression is perhaps one of the

most well-known and well understood

algorithms in statistics and machine

learning. Linear regression is a linear

model. For example, a model that assumes

a linear relationship between the input

variables x and the single output

variable y. And you'll see this if you

remember from your algebra classes, y =

mx + c. Imagine we are predicting

distance traveled y from speed x. Our

linear regression model representation

for this problem would be y = m * x + c

or distance = m * speed + c where m is

the coefficient and c is the y

intercept. And we're going to look at

two different variations of this. First,

we're going to start with time is

constant. And you can see we have a

bicyclist. He's got a safety gear on,

thank goodness. Speed equals 10

meters/s. And so over a certain amount

of time, his distance equals 36 km. We

have a second bicyclist who's going

twice the speed or 20 m/s. And you can

guess if he's going twice the speed and

time is a constant, then he's going to

go twice the distance. And that's easy

to compute. 36 * 2, you get 72 km. And

so if you had the question of how fast

would somebody going three times that

speed or 30 m/s is, you can easily

compute the distance in our head. We can

do that without needing a computer, but

we want to do this for more complicated

data. So, it's kind of nice to compare

the two. But, let's just take a look at

that and what that looks like in a

graph. So, in a linear regression model,

we have our distance to the speed and we

have our m equals the ve slope of the

line. And we'll notice that the line has

a plus slope. And as the speed

increases, distance also increases.

Hence, the variables have a positive

relationship. And so your speed of the

person which equals y = mx plus c

distance traveled in a fixed interval of

time. And we could very easily compute

either following the line or just

knowing it's 3 * 10 m/s that this is

roughly 102 km distance that this third

bicycle has traveled. One of the key

definitions on here is positive

relationship. So the slope of the line

is positive. As distance increase so

does speed increase. Let's take a look

at our second example where we put

distance is a constant. So we have speed

equals 10 m/s. They have a certain

distance to go and it takes him 100

seconds to travel that distance. And we

have our second bicyclist who's still

doing 20 m/s. Since he's going twice the

speed, we can guess he'll cover the

distance in about half the time, 50

seconds. And of course, you could

probably guess on the third one, 100

divided by 30 since he's going three

times the speed. You can easily guess

that this is 33.3333

seconds time. We put that into a linear

regression model or a graph. If the

distance is assumed to be constant,

let's see the relationship between speed

and time. And as time goes up, the

amount of speed to go that same distance

goes down. So now your m equals a minus

v slope of the line. As the speed

increases, time decreases. Hence, the

variable has a negative relationship.

Again, there's our definition. positive

relationship and negative relationship

dependent on the slope of the line and

with a simple formula like this um and

even a significant amount of data. Let's

uh see what the mathematical

implementation of linear regression and

we'll take this data. So suppose we have

this data set where we have xyx= 1 2 3 4

5 standard series and the y value is 3

22 43. When we take that and we go ahead

and plot these points on a graph, you

can see there's kind of a nice

scattering and you could probably

eyeball a line through the middle of it.

But we're going to calculate that exact

line for linear regression. And the

first thing we do is we come up here and

we have the mean of Xi. And remember

mean is basically the average. So we

added five plus 4 plus 3 plus 2 plus 1

and divide by five. And that simply

comes out as three. And then we'll do

the same for y. We'll go ahead and add

up all those numbers and divide by five.

And we end up with a mean value of y of

i equals 2.8 where the x i references

it's an average or means value. And the

yi also equals a means value of y. And

when we plot that, you'll see that we

can put in the y= 2.8 and the x= 3 in

there on our graph. We kind of gave it a

little different color so you could sort

it out with the dashed lines on it. And

it's important to note that when we do

the linear regression, the linear

regression model should go through that

dot. Now, let's find our regression

equation to find the best fit line.

Remember, we go ahead and take our y= mx

plus c. So, we're looking for m and c.

So, to find this equation for our data,

we need to find our slope of m and our

coefficient of c. And we have y = mx + c

where m equals the sum of x - x average

* y - y average or y means and x means

over the sum of x - x means squared.

That's how we get the slope of the value

of the line. And we can easily do that

by creating some columns here. We have

xy. Computers are really good about

iterating through data. And so we can

easily compute this and fill in a graph

of data. And in our graph you can easily

see that if we have our x value of 1 and

if you remember the x i or the means

value is 3. 1 - 3 equals a -2 and 2 - 3

= a -1 so on and so forth. And we can

easily fill in the column of x - x i y -

yi. And then from those we can compute x

- x i^ 2 and x - x i * y - yi. And you

can guess it that the next step is to go

ahead and sum the different columns for

the answers we need. So we get a total

of 10 for our x - x i^2 and a total of 2

for x - x i * y - yi. And we plug those

in, we get 2/10, which equals2. So now

we know the slope of our line equals2.

So we can calculate the value of c.

That'd be the next step is we need to

know where it crosses the y ais. And if

you remember, I mentioned earlier that

the linear regression line has to pass

through the means value, the one that we

showed earlier. We can just flip back up

there to that graph. And you can see

right here, there's our means value,

which is 3 x= 3 and y= 2.8. And since we

know that value, we can simply plug that

into our formula. Y =2x + c. So we plug

that in, we get 2.8 8 =2 * 3 + C. And

you can just solve for C. So now we know

that our coefficient equals 2.2. And

once we have all that, we can go ahead

and plot our regression line. Y =2 * X +

2.2. And then from this equation, we can

compute new values. So let's predict the

values of Y using X= 1 2 3 4 5 and plot

the points. Remember the 1 2 3 4 5 was

our original x values. So now we're

going to see what y thinks they are, not

what they actually are. And we plug

those in, we get y of designated with y

of p. You can see that x= 1 = 2.4, x= 2=

2.6, and so on and so on. So we have our

y predicted values of what we think it's

going to be when we plug those numbers

in. And when we plot the predicted

values along with the actual values, we

can see the difference. And this is one

of the things that's very important with

linear regression in any of these models

is to understand the error. And so we

can calculate the error on all of our

different values. And you can see over

here we plotted um x and y and y

predict. And we draw a little line so

you can sort of see what the error looks

like there between the different points.

So our goal is to reduce this error. We

want to minimize that error value on our

linear regression model. Minimizing the

distance. There are lots of ways to

minimize the distance between the line

and the data points like sum of squared

errors, sum of absolute errors, root

mean square error, etc. We keep moving

this line through the data points to

make sure the best fit line has the

least squared distance between the data

points and the regression line. So to

recap with a very simple linear

regression model, we first figure out

the formula of our line through the

middle and then we slowly adjust the

line to minimize the error. Keep in mind

this is a very simple formula. The math

gets even though the math is very much

the same, it gets much more complex as

we add in different dimensions. So this

is only two dimensions. Y equals MX + C.

But you can take that out to X ZQ all

the different features in there and they

can plot a linear regression model on

all of those using the different

formulas to minimize the error. Let's go

ahead and take a look at decision trees.

A very different way to solve problems

in the linear regression model. Decision

tree is a treeshaped algorithm used to

determine a course of action. Each

branch of a tree represents a possible

decision, occurrence, or reaction. We

have data which tells us if it is a good

day to play golf. And if we were to open

this data up in a general spreadsheet,

you can see we have the outlook, whether

it's rainy, overcast, sunny,

temperature, hot, mild, cool, humidity,

windy, and did I like to play golf that

day? Yes or no. So, we're taking a

census. And certainly, I wouldn't want a

computer telling me when I should go

play golf or not. But you could imagine

if you got up in the night before,

you're trying to plan your day and it

comes up and says, "Tomorrow would be a

good day for golf for you in the morning

and not a good day in the afternoon or

something like that." This becomes very

beneficial and we see this in a lot of

applications coming out now where it

gives you suggestions and lets you know

what what would uh fit the match for you

for the next day or the next purchase or

the next uh whatever you know next mail

out in this case is tomorrow a good day

for playing golf based on the weather

coming in. And so we come up and let's

uh determine if you should play golf

when the day is sunny and windy. So we

found out the forecast tomorrow is going

to be sunny and windy. And suppose we

draw our tree like this. We're going to

have our humidity. And then we have our

normal, which is uh if it's if you have

a normal humidity, you're going to go

play golf. And if the humidity is really

high, then we look at the outlook. And

if the outlook is sunny, overcast, or

rainy, it's going to change what you

choose to do. So if you know that it's a

very high humidity and it's sunny,

you're probably not going to play golf

cuz you're going to be out there

miserable, fighting off the mosquitoes

that are out joining you to play golf

with you. Maybe if it's rainy, you

probably don't want to play in the rain.

But if it's slightly overcast and you

get just the right shadow, that's a good

day to play golf and be outside out on

the green. Now, in this example, you can

probably make your own tree pretty

easily cuz it's a very simple set of

data going in. But the question is, how

do you know what to split? Where do you

split your data? What if this is much

more complicated data where it's not

something that you would particularly

understand? like studying cancer, they

take about 36 measurements of the

cancerous cells and then each one of

those measurements represents how

bulbous it is, how extended it is, how

sharp the edges are, something that as a

human we would have no understanding of.

So how do we decide how to split that

data up and is that the right decision

tree? But so that's a question that's

going to come up. Is this the right

decision tree? For that we should

calculate entropy and information gain.

Two important vocabulary words there are

the entropy and the information gain.

Entropy. Entropy is a measure of

randomness or impurity in the data set.

Entropy should be low. So we want the

chaos to be as low as possible. We don't

want to look at it and be confused by

the images or what's going on there with

mixed data. And the information gain, it

is a measure of decrease in entropy

after the data set is split. Also known

as entropy reduction. information gain

should be high. So we want our

information that we get out of the split

to be as high as possible. Let's take a

look at entropy from the mathematical

side. In this case, we're going to

denote entropy as I of P of and N where

P is the probability that you're going

to play a game of golf and N is the

probability where you're not going to

play the game of golf. Now, you don't

really have to memorize these formulas.

There's a few of them out there

depending on what you're working with.

But it's important to note that this is

where this formula is coming from. So

when you see it, you're not lost when

you're running your programming, unless

you're building your own decision tree

code in the back. And we simply have a

log 2 of p + n minus n / p + n * the log

squar of n of p plus n. But let's break

that down and see what actually looks

like when we're computing that from the

computer script side. Entropy of a

target class of the data set is the

whole entropy. So we have entropy play

golf. And we look at this. If we go back

to the data, you can simply count how

many yeses and no in our complete data

set for playing golf days. In our

complete set, we find we have five days

we did play golf and nine days we did

not play golf. And so our I equals, if

you add those together, 9 + 5 is 14. And

so our I equals 5 over 14 and 9 over 14.

That's our PNN values that we plug into

that formula. And you can go 5 over

14=.36.

9 over4=64.

And when you do the whole equation, you

get the -.36

log^ 2 of.36 minus.64 log of

64. And we get a set value. We get 94.

So we now have a full entropy value for

the whole set of data that we're working

with. And we want to make that entropy

go down. And just like we calculated the

entropy out for the whole set, we can

also calculate entropy for playing golf

and the outlook. Is it going to be

overcast or rainy or sunny? And so we

look at the entropy. We have P of sunny

times E of three of two. And that just

comes out how many sunny days yes and

how many sunny days no over the total,

which is five. Don't forget to put the

we'll divide that five out later on.

equals P overcast = 4 comma 0 plus rainy

= 2a 3 and then when you do the whole

setup we have 5 over4 remember I said

there was a total of five 5 over 14 *

the i of 3 of 2 + 4 over 14 * the 4 0

and 514 over i of 23 and so we can now

compute the entropy of just the part

that has to do with the forecast and we

get 693 similar We can calculate the

entropy of other predictors like

temperature, humidity and wind. And so

we look at the gain outlook. How much

are we going to gain from this entropy

play golf minus entropy play golf

outlook? And we can take the original

0.94 for the whole set minus the entropy

of just the rainy day and temperature

and we end up with a gain of.247.

So this is our information gain.

Remember we define entropy and we define

information gain. The higher the

information gain, the lower the entropy,

the better. The information gain of the

other three attributes can be calculated

in the same way. So we have our gain for

temperature equals 0.029.

We have our gain for humidity

equals.152.

And our gain for a windy day equals

0048. And if you do a quick comparison,

you'll see the 247 is the greatest gain

of information. So that's the split we

want. Now let's build the decision tree.

So, we have the outlook. Is it going to

be sunny, overcast, or rainy? That's our

first split because that gives us the

most information gain. And we can

continue to go down the tree using the

different information gains with the

largest information. We can continue

down the nodes of the tree where we

choose the attribute with the largest

information gain as the root node and

then continue to split each subnode with

the largest information gain that we can

compute. And although it's a little bit

of a tongue twister to say all that, you

can see that it's a very easy to view

visual model. We have our outlook. We

split it three different directions. If

the outlook is overcast, we're going to

play. And then we can split those

further down if we want. So if the over

outlook is sunny, but then it's also

windy. If it's uh windy, we're not going

to play. If it's uh not windy, we'll

play. So, we can easily build a nice

decision tree to guess what we would

like to do tomorrow and give us a nice

recommendation for the day. So, we want

to know if it's a good day to play golf

when it's sunny and windy. Remember the

original question that came out,

tomorrow's weather report is sunny and

windy. You can see by going down the

tree, we go outlook sunny, outlook

windy. We're not going to play golf

tomorrow. So, our little smartwatch pops

up and says, I'm sorry, tomorrow's not a

good day for golf. It's going to be

sunny and windy. And if you're a huge

golf fan, you might go, "Uh oh, it's not

a good day to play golf." We can go in

and watch a golf game at home. So, we'll

sit in front of the TV instead of being

out playing golf in the wind. Now that

we looked at our decision tree, let's

look at the third one of our algorithms

we're investigating. Support vector

machine. Support vector machine is a

widely used classification algorithm.

The idea of support vector machine is

simple. The algorithm creates a

separation line which divides the

classes in the best possible manner. For

example, dog or cat, disease or no

disease. Suppose we have a labeled

sample data which tells height and

weight of males and females. A new data

point arrives and we want to know

whether it's going to be a male or a

female. So we start by drawing a line.

We draw decision lines. But if we

consider decision line one, then we will

classify the individual as a male. And

if we consider decision line two, then

it'll be a female. So you can see this

person kind of lies in the middle of the

two groups. So it's a little confusing

trying to figure out which line they

should be under. We need to know which

line divides the classes correctly. But

how the goal is to choose a hyper plane

and that is one of the key words they

use when we talk about support vector

machines. Choose a hyper plane with the

greatest possible margin between the

decision line and the nearest point

within the training set. So you can see

here we have our support vector. We have

the two nearest points to it and we draw

a line between those two points. And the

distance margin is the distance between

the hyper plane and the nearest data

point from either set. So we actually

have a value and it should be equal

distant between the two points that

we're comparing it to. When we draw the

hyperplanes, we observe that line one

has a maximum distance. So we observe

that line one has a maximum distance

margin. So we'll classify the new data

point correctly. And our result on this

one is going to be that the new data

point is MEL. One of the reasons we call

it a hyper plane versus a line is that a

lot of times we're not looking at just

weight and height. We might be looking

at 36 different features or dimensions.

And so when we cut it with a hyper

plane, it's more of a three-dimensional

cut in the data, multi-dimensional that

cuts the data a certain way. And each

plane continues to cut it down until we

get the best fit or match. Let's

understand this with the help of an

example. Problem statement. You always

start with a problem statement when

you're going to put some code together.

We're going to do some coding now.

Classifying muffin and cupcake recipes

using support vector machines. So the

cupcake versus the muffin. Let's have a

look at our data set. And we have the

different recipes here. We have a muffin

recipe that has so much flour. I'm not

sure what measurement 55 is in, but it

has 55, maybe it's ounces, but it has a

certain amount of flour, certain amount

of milk, sugar, butter, egg, baking

powder, vanilla, and salt. And so based

on these measurements, we want to guess

whether we're making a muffin or a

cupcake. And you can see in this one, we

don't have just two features. We don't

just have height and weight as we did

before between the male and female. In

here, we have a number of features. In

fact, in this, we're looking at eight

different features to guess whether it's

a muffin or a cupcake. What's the

difference between a muffin and a

cupcake? Turns out muffins have more

flour, while cupcakes have more butter

and sugar. So, basically, the cupcakes a

little bit more of a dessert, where the

muffin's a little bit more of a fancy

bread. But how do we do that in Python?

How do we code that to go through

recipes and figure out what the recipe

is? And I really just want to say

cupcakes versus muffins like some big

professional wrestling thing. Before we

start in our cupcakes versus muffins, we

are going to be working in Python.

There's many versions of Python, many

different editors. That is one of the

strengths and weaknesses of Python is it

just has so much stuff attached to it.

It's one of the more popular data

science programming packages you can

use. In this case, we're going to go

ahead and use Anaconda in Jupyter

Notebook. The Anaconda Navigator has all

kinds of fun tools. Once you're into the

Anaconda Navigator, you can change

environments. I actually have a number

of environments on here. We'll be using

Python 36 environment. So, this is in

Python version 36. Although, it doesn't

matter too much which version you use. I

usually try to stay with the 3x because

they're current unless you have a

project that's very specifically in

version 2x 27 I think is usually what

most people use in the version two. And

then once we're in our um Jupiter

notebook editor, I can go up and create

a new file and we'll just jump in here.

In this case, we're doing SPM muffin

versus cupcake. And then let's start

with our packages for data analysis.

And we almost always use a couple

there's a few very standard packages we

use. We use import oops import

numpy

that's for number python. They usually

denote it as np that's very comma that's

very common. And then we're going to

import pandas as pd. And numpy deals

with number arrays. There's a lot of

cool things you can do with the numpy uh

setup as far as multiplying all the

values in an array in a numpy array data

array. Pandas I can't remember if we're

using it actually in this data set. I

think we do as an import it makes a nice

data frame. And the difference between a

data frame and a numpy array is that a

data frame is more like your Excel

spreadsheet. You have columns, you have

indexes. So you have different ways of

referencing it easily viewing it. And

there's additional features you can run

on a data frame. And pandas kind of sits

on numpy. So they you need them both in

there. And then finally, we're working

with the support vector machine. So from

sklearn, we're going to use the sklearn

model. Import SVM support vector

machine.

And then as a data scientist, you should

always try to visualize your data. Some

data obviously is too complicated or

doesn't make any sense to the human. But

if it's possible, it's good to take a

second look at it so that you can

actually see what you're doing. Now, for

that, we're going to use two packages.

We're going to import mapplot

library.pipplot as plt. Again, very

common. And we're going to import seabor

as sns. And we'll go ahead and set the

font scale in the SNS right in our

import line. That's what this U

semicolon followed by a line of data.

We're going to set the SNS. And these

are great because the the seabour sits

on top of map plot library just like

pandas sits on numpy. So it adds a lot

more features and uses and control.

We're obviously not going to get into

mattplot library and seabour. It' be its

own tutorial. We're really just focusing

on the SVM, the support vector machine

from sklearn. And since we're in Jupyter

notebook, uh we have to add a special

line in here for our mattplot library.

And that's your percentage sign or amber

sign mattplot library in line. Now, if

you're doing this in just a straight

code project, a lot of times I use like

Notepad++

and I'll run it from there. You don't

have to have that line in there because

it'll just pop up as its own window on

your computer depending on how your

computer's set up because we're running

this in the Jupyter notebook as a

browser setup. This tells it to display

all of our graphics right below on the

page. So that's what that line is for.

Remember the first time I ran this, I

didn't know that and I had to go look

that up years ago. It's quite a

headache. So mattplot library inline is

just because we're running this on the

web setup and we can go ahead and run

this. make sure all our modules are in.

They're all imported, which is great. If

you don't have them import, you'll need

to go ahead and pip. Use the pip or

however you do it. There's a lot of

other install packages out there,

although pip is the most common. And you

have to make sure these are all

installed on your Python setup. The next

step, of course, is we got to look at

the data. You can't run a model for

predicting data if you don't have actual

data. So, to do that, let me go ahead

and open this up and take a look. And we

have our uh cupcakes versus muffins. and

it's a CSV file or CSV meaning that it's

commaepparated variable

and it's going to open it up in a nice

uh spreadsheet for me. And you can see

up here we have the type we have muffin

muffin muffin cupcake cupcake cupcake

and then it's broken up into flour,

milk, sugar, butter, egg, baking powder,

vanilla and salt. So we can do is we can

go ahead and look at this data also in

our Python.

Let us create a variable recipes equals

we're going to use our pandas module

read CSV. Remember is a commaepparated

variable

and the file name happened to be

cupcakes versus muffins. Oops, I got

double brackets there.

Do it this way.

There we go. cupcakes versus muffins.

Because the program I loaded or the the

place I saved this particular Python

program is in the same folder, we can

get by with just the file name. But

remember, if you're storing it in a

different location, you have to also put

down the full path on there.

And then because we're in pandas, we're

going to go ahead and you can actually

in line you can do this, but let me do

the full print. You can just type in

recipes.head head in the Jupyter

notebook. But if you're running in code

in a different script, you'd need to go

ahead and type out the whole print

recipes.

And Pandanda's knows that's going to do

the first five lines of data. And if we

flip back on over to the spreadsheet

where we opened up our CSV file,

uh you can see where it starts on line

two. This one calls it zero. And then 2

3 4 5 6 is going to match. Go and close

that out because we don't need that

anymore. And it always starts at zero.

And these are it automatically indexes

it since we didn't tell it to use an

index in here. So that's the index

number for the left hand side. And it

automatically took the top row as

labels. So pandas using it to read a CSV

is just really slick and fast. One of

the reasons we love our pandas, not just

because they're cute and cuddly teddy

bears.

And let's go ahead and plot our data.

And I'm not going to plot all of it. I'm

just going to plot the uh sugar and

flour. Now, obviously, you can see where

they get really complicated if we have

tons of different features. And so,

you'll break them up and maybe look at

just two of them at a time to see how

they connect.

And to plot them, we're going to go

ahead and use Seabor. So, that's our

SNS. And the command for that is SNS.LM

plot. And then the two different

variables I'm going to plot is flour and

sugar.

Data equals recipes. The hue equals

type. And this is a lot of fun because

it knows that this is pandas coming in.

So this is one of the powerful things

about pandas mixed with seabor and doing

graphing. And then we're going to use a

pallet set one. There's a lot of

different sets in there. You can go look

them up for seabor. We do a regular fit

regular equals false. So, we're not

really trying to fit anything. And it's

a scatter KWS.

A lot of these settings you can look up

in Seabor. Half of these you could

probably leave off when you run them.

Somebody played with this and found out

that these were the best settings for

doing a Seabor plot. And let's go ahead

and run that. And because it does it in

line, it just puts it right on the page.

And you can see right here that just

based on sugar and flour alone, there's

a definite split. And we use these

models because you can actually look at

it and say, "Hey, if I drew a line right

between the middle of the blue dots and

the red dots, we'd be able to do an SVM

and and a hyper plane right there in the

middle.

Then the next step is to format or

pre-process

our data.

And we're going to break that up into

two parts.

We need a type label. And remember,

we're going to decide whether it's a

muffin or a cupcake. Well, a computer

doesn't know muffin or cupcake. It knows

zero and one. So, what we're going to do

is we're going to create a type label.

And from this we'll create a numpy array

nump where and this is where we can do

some logic. We take our recipes from our

panda and wherever type equals muffin

it's going to be zero. And then if it

doesn't equal muffin which is cupcakes

it's going to be one. So we create our

type label. This is the answer. So when

we're doing our training model remember

we have to have a a training data. This

is what we're going to train it with. Is

that it's zero or one? it's a muffin or

it's not.

And then we're going to create our

recipe features.

And if you remember correctly from right

up here, the first column is type.

So we really don't need the type column

because that's our muffin or cupcake.

And in pandas, we can easily sort that

out.

We take our value recipes

columns. That's a pandas function built

into pandas.

values converting them to values. So

it's just the column titles going across

the top and we don't want the first one.

So what we do is since it always starts

at zero, we want one

colon till the end.

And then we want to go ahead and make

this a list. And this converts it to a

list of strings.

And then we can go ahead and just take a

look and see what we're looking at for

the features. Make sure it looks right.

Me go ahead and run that.

And I forgot the S on recipes. So, we'll

go ahead and add the S in there and then

run that. And we can see we have flour,

milk, sugar, butter, egg, baking powder,

vanilla, and salt. And that matches what

we have up here, right? Where we printed

out everything but the type. So, we have

our features and we have our label.

Now, the recipe features is just the

titles of the columns. We actually need

the ingredients.

And at this point, we have a couple

options. One, we could run it over all

the ingredients.

And when you're doing this, usually you

do. But for our example, we want to

limit it so you can easily see what's

going on because if we did all the

ingredients, we have, you know, that's

what, um, seven, eight different

hyperplanes that would be built into it.

We only want to look at one. So you can

see what the SVM is doing.

And so we'll take our recipes and we'll

do just flour and sugar. Again, you can

replace that with your recipe features

and do all of them, but we're going to

do just flour and sugar. And we're going

to convert that to values. We don't need

to make a list out of it because it's

not string values. These are actual

values on there. And we can go ahead and

just print

ingredients. And you can see what that

looks like.

Uh, and so we have just the nanoflower

and sugar, just the two sets of plots.

And just for fun, let's go ahead and

take this over here and take our recipe

features.

And so if we decided to use all the

recipe features, you'll see that it

makes a nice column of different data.

So it just strips out all the labels and

everything. We just have just the

values. But because we want to be able

to view this easily in a plot later on,

we'll go ahead and take that and just do

flour and sugar.

And we'll run that. And you'll see it's

just the two columns.

So the next step is to go ahead and fit

our model.

We'll go ahead and just call it model.

And it's a SVM. We're using a package

called SVC.

In this case, we're going to go ahead

and set the kernel equals linear. So,

it's using a specific setup on there.

And if we go to the reference on their

website for the SVM,

you'll see that there's about there's

eight of them here. Three of them are

for regression.

Three are for classification. The SVC,

support vector classification, is

probably one of the most commonly used.

And then there's also one for detecting

outliers and another one that has to do

with something a little bit more

specific on the model. But SVC and SVR

are the two most commonly used standing

for support vector classifier and

support vector regression. Remember

regression is an actual value, a float

value or whatever you're trying to work

on. And SBC is a classifier. So it's a

yes, no, true, false.

But for this we want to know 01 muffin

cupcake. If we go ahead and create our

model and once we have our model

created, we're going to do model.fit.

And this is very common, especially in

the sklearn. All their models are

followed with the fit command.

And what we put into the fit, what we're

training with it is we're putting in the

ingredients, which in this case we

limited to just flour and sugar, and the

type label. Is it a muffin or cupcake?

Now, in more complicated data science

series, you'd want to split into, we

won't get into that today, where you

split it into training data and test

data. And they even do something where

they split it into thirds, where a third

is used for where you switch between

which one's training and test. There's

all kinds of things go into that. It

gets very complicated when you get to

the higher end. Not overly complicated,

just an extra step, which we're not

going to do today because this is a very

simple set of data.

And let's go ahead and run this. And now

we have our model fit. And uh I got an

error here. So let me fix that real

quick. It's capital SBC. It turns out

I did it lowercase.

Support vector

classifier. There we go. Let's go ahead

and run that. And you'll see it comes up

with all this information that it prints

out automatically. These are the

defaults of the model. You notice that

we changed the kernel to linear. And

there's our kernel linear on the

printout. And there's other different

settings you can mess with.

We're going to just to leave that alone

for right now. For this, we don't really

need to mess with any of those.

So, next we're going to dig a little bit

into our newly trained model. And we're

going to do this so we can show you on a

graph.

And let's go ahead and get the

separating.

and we're going to say uh we're going to

use a W for our variable on here and

we're going to do model.coreeficient_0.

So what the heck is that? Again, we're

digging into the model. So we've already

got a prediction and a train. This is a

math behind it that we're looking at

right now. And so the w is going to

represent two different coefficients.

And if you remember, we had y = mx + c.

So these coefficients are connected to

that but in two-dimensional it's a

plane.

We don't want to spend too much time on

this because you can get lost in the

confusion of the math. So if you're a

math wiz this is great. You can go

through here and you'll see that we have

a= minus w of 0 over w of 1. Remember

there's two different values there. And

that's basically the slope that we're

generating.

And then we're going to build an xx.

What is xx? We're going to set it up to

a numpy array. There's our np line

space. So we're creating a line

of values between 30 and 60. So it just

creates a set of numbers for x. And then

if you remember correctly, we have our

formula y equals the slope * x

plus the intercept. Well, to make this

work, we can do this as y

equals the slope times each value in

that array. That's the neat thing about

numpy. So, when I do a * xx, which is a

whole numpy array of values, it

multiplies a across all of them. And

then it takes those same values and we

subtract the model intercept. That's

your uh we had mx plus c. So, that'd be

the c from the formula y mx plus c.

And that's where all these numbers come

from. A little bit confusing because

it's digging out of these different

arrays. And then what we want to do is

we're going to take this and we're going

to go ahead and plot it. So plot the

parallels to separating hyper plane that

pass through the support vectors. And so

we're going to create B equals a model

support vectors. Pulling our support

vectors out there. Here's our y, which

we now know is a set of data. And we

have uh we're going to create y down = a

* xx + b1 - a * b 0. And then model

support vector b is going to be set that

to a new value the minus1 setup. And y y

up = a * xx + b1 - a * b 0. And we can

go ahead and just run this to load these

variables up. If you wanted to know

understand a little bit more of what's

going on, you can see if we print

y, let me just run that. You can see

it's an array. This is a line. It's

going to have in this case between 30

and 60. So there's going to be 30

variables in here. And the same thing

with y y up y y y y y y y y y y y y y y

y y y y y y y y y y y y y y y y y y y y

y y y y y y y y y y y y y y y y y y y y

y y y y y y y y y y y y y y y y y y y y

y y y y y y down and we'll we'll plot

those in just a minute on a graph so you

can see what those look like.

Just go ahead and delete that out of

here and run that. So, it loads up the

variables. Nice clean slate. I'm just

going to copy this from before. Remember

this? Our SNS, our Seabor plot, LM plot,

flower, sugar. And I'll just go and run

that real quick so you can see what

remember what that looks like. It's just

a straight graph on there. And then one

of the neat things is because Seabour

sits on top of piplot,

we can do the piplot for the line going

through. And that is simply plt.plot

And that's our xx and y are two

corresponding values xy. And then

somebody played with this to figure out

that the line width equals 2 and the

color black would look nice. So let's go

ahead and run this whole thing with the

pi plot on there. And you can see when

we do this, it's just doing flour and

sugar on here.

Corresponding line between the sugar and

the flour and the muffin versus cupcake.

Um, and then we generated the support

vectors, the y down and y up. So let's

take a look and see what that looks

like.

So we'll do our plot.

And again, this is all against xx,

our x value, but this time we have y

down.

And let's do something a little fun with

this. We can put in a k dash dash. That

just tells it to make it a dotted line.

And if we're going to do the down one,

we also want to do the up one. So here's

our y

up. And when we run that, it adds both

sets of line. And so here's our support.

And this is what you expect. You expect

these two lines to go through the

nearest data point. So the dash lines go

through the nearest muffin and the

nearest cupcake when it's plotting it.

And then your SVM goes right down the

middle. So it gives it a nice split in

our data. And you can see how easy it is

to see based just on sugar and flour

which one's a muffin or a cupcake.

Let's go ahead and create a function

to predict

muffin or cupcake.

I've got my uh recipes. I pulled off the

um internet and I want to see the

difference between a

muffin or a cupcake. And so we need a

function to push that through. And uh we

create a function with deaf. And let's

call it muffin or cupcake. And remember,

we're just doing flour and sugar today.

We're not doing all the ingredients. And

that actually is a pretty good split.

You really don't need all the

ingredients to know it's flour and

sugar. And let's go ahead and do an if

else statement. So if model predict

is of flower and sugar equals zero. So

we take our model and we do run a

predict. It's very common in sklearn

where you have a predict. You put the

data in and it's going to return a

value. In this case if it equals zero

then print you're looking at a muffin

recipe. Else if it's not zero that means

it's one and you're looking at a cupcake

recipe. That's pretty straightforward

for

function or def for definition. Deaf is

how you do that in Python. And of

course, if you're going to create a

function, you should run something in

it. And so, let's run a cupcake. And

we're going to send it values 50 and 20.

A muffin or a cupcake. I don't know what

it is. And let's run this and just see

what it gives us. It says, "Oh, it's a

muffin. You're looking at a muffin

recipe." So, it very easily predicts

whether we're looking at a muffin or a

cupcake recipe. Let's plot this. There

we go. Plot this on the graph so we can

see what that actually looks like. And

I'm just going to copy and paste it from

below where we plotting all the points

in there.

So, this is nothing different than we

did before. If I run it, you'll see it

has all the points and the lines on

there. And what we want to do is we want

to add another point. And we'll do

pltot.

And if you remember correctly, we did

for our test we did 50

and 20. And then somebody went in here

and decided we'll do yo for yellow or

it's kind of a orangeish yellow color is

going to come out. Marker size nine.

Those are settings you can play with.

Somebody else played with them to come

up with the right setup so it looks

good. And you can see there it is

graphed clearly a muffin.

In this case in cupcakes versus muffins,

the muffin has won. And if you'd like to

do your own muffin cupcake contender

series, you certainly can send a note

down below and the team at SimplyLearn

will send you over the data they use for

the muffin and cupcake. And that's true

of any of the data. We didn't actually

run a plot on it earlier. We had men

versus women. You can also request that

information to run it on your data

setup. So you can test that out.

So to go back over our setup, we went

ahead for our support vector machine

code. We did a predict 40 parts flour,

20 parts sugar. I think it was different

than the one we did whether it's a

muffin or a cupcake. Hence, we have

built a classifier using SVM which is

able to classify if a recipe is of a

cupcake or a muffin. Which wraps up our

cupcake versus muffin. So the key

takeaways, what is machine learning? We

discussed that with some of the

different aspects of machine learning on

there. We went into types of machine

learning. If you memorize we have

supervised, unsupervised and

reinforcement learning. We discussed

regression line or best fit and we did

the building a decision tree and what

the logic is behind that. And finally we

did classification using SVM support

vector machine and we did the code in

there. Today we are diving into machine

learning, the technology behind things

like Netflix recommendations, CD, and

even the face unlock of your phone.

Machine learning helps devices get

smarter by learning from data and

predicting what we might like or need.

And here's why machine learning is huge

for your career. Right now, machine

learning jobs are among the fastest

growing roles worldwide. Companies in

every industry, tech, healthcare,

finance, and more, are looking for

people with machine learning skills to

improve their products, automate tasks,

and make smarter decisions. Machine

learning engineers in the US earn around

$112,000 on average with plenty of room

for growth as you gain experience. So,

if you want to jump into this exciting

field, learning machine learning can

open doors to highpaying in- demand

jobs. So in this video I'll guide you

through the ultimate road map to master

machine learning in 2025 one step at a

time. So let's get started. So in the

first month start with the foundations

of programming. So programming is a

language you'll use to communicate with

your computer and bring machine learning

algorithms to life. So this month is all

about Python, the language of choice for

most machine learning practitioners. So

here's what to focus on. First, learn

Python basics. Begin with Python's

fundamentals like variables, data types,

loops and functions. So spend time

writing small programs daily to get

comfortable. After that explore the key

libraries like numpy, pandas and

scikitlearn. So numpy is for numerical

operations. It makes handling large data

sets faster and easier. And pandas is to

manipulate and analyze data. So pandas

allow you to filter, sort and reshape

data in a breeze. And then scikitlearn

is for implementing algorithms in just a

few lines of code. So now you might have

heard about R, another language used in

machine learning. But don't stress about

it now. Python will serve you well,

especially as a beginner, because it's

simpler and more flexible. So aim to

spend an hour or two each day coding. By

the end of this month, you'll have a

solid base to build on. Now, in the

second month, get organized with version

control and data structures. So this

month is about learning how to organize

and manage your code effectively and

sharpening your problem solving skills

with data structures and algorithms. So

first is version control with git. So

think of git as your project history

tracker. So imagine working on a big

project and making changes then

realizing something went wrong. You want

to go back to an earlier version, right?

So that's where git comes in. And here's

what you should practice. Number one is

committing changes. So save different

versions of your work as you progress.

And then branching which means work on

separate features without affecting your

main code. And then comes merging which

means combining changes from different

versions once they are ready. So you

have to set up an account on GitHub or

GitLab to store your projects online. So

not only will this be super useful, but

it'll also start building your

portfolio. Now next is data structures

and algorithm. So think of data

structures like tools in a toolkit. So

each one like arrays, stacks, cues, etc.

serves a specific purpose. So here's how

to approach them. Number one, arrays and

lists. Now arrays and lists are for

storing data in sequence. After that,

you can get familiar with stacks and

cues. So stacks and cues are for tasks

that need ordered data access. And then

you have sorting and searching

algorithms. So these make your programs

more efficient. And that's super

important in machine learning where data

can get massive. So the goal here is to

build up your problem solving skills

which are key to machine learning

success. So take it slow, practice daily

and you'll see progress. Now in the

third month, learn to access data with

SQL. So in machine learning, a lot of

work involves accessing and organizing

data from databases. So SQL, a

structured query language, is your

ticket to getting the data you need for

training ML models. So here's what you

should focus on. Select and where. So

these commands help you pull specific

pieces of data and then you can move on

to joins. Joins usually combine data

from different tables. So this is so

powerful that you'll use it all the

time. And then comes group by and

aggregate functions. They are great for

summarizing data to find patterns. So

spend time working with sample databases

you can find online and practice writing

queries. Being comfortable with SQL will

save you time when preparing data for

your models. Now after completing the

third month you can move on to

mathematics which is building your

analytical mind. So this month we are

tackling the math behind machine

learning. So don't worry you don't need

to be a math genius but understanding

certain concepts will make everything

feel less mysterious. So in this month

you have to focus on linear algebra. So

this is the math behind how models see

data. So you can study vectors, matrices

and operations like multiplication. Next

comes calculus. So you'll use calculus

to help your models learn. So you have

to focus on derivatives and gradients

which help minimize errors in your

model. And then you can move on to

probability and statistics. So

understanding probability helps you make

sense of data. So learn about

distributions like normal distribution,

bormal distribution and then variance

and standard deviation. So once you have

learned maths, next you'll be moving on

to data handling and visualization which

is the heart of machine learning as you

all know. So with Matt under your belt,

it's time to dig into data handling and

visualization. So data preparation is

vital because your model is only as good

as the data you feed it. So number one

comes data manipulation. So using pandas

and numpy, you'll clean and organize

your data. You might be removing missing

values like clean up messy data so it

doesn't confuse your model. And then

you'll learn transforming variables like

converting data into formats that work

for models. And then you will move on to

encoding categorical data like changing

text data like female or male into

numbers. Now once you're done with data

manipulation, next comes data

visualization. So visualization is how

you get to see your data before training

a model. So here you have to learn

mattplot lip and seabboard. So you can

create line charts, histograms, scatter

plots and heat maps. So this lets you

explore patterns and spot outliers. So

understanding these patterns in your

data is crucial for building effective

models. Now in the sixth month you'll be

moving on to the machine learning

fundamentals. So now it's time to start

building your own models. So you will

focus on two main types of machine

learning this month. Number one comes

the supervised learning. So this is when

you train a model on label data where

the outcome is already known. So you'll

work with algorithms like linear

regression which predicts a continuous

outcome. Then you'll work with decision

trees which breaks down decisions into a

tree structure. And then you have

support vector machines under supervised

learning which updates data into

classes. Now after supervised learning

comes unsupervised learning. So here

your model identifies patterns in data

without labeled outcomes. So two popular

techniques in unsupervised learning is

number one clustering like K means

clustering which means group similar

data points and then you have

dimensionality reduction. This reduces

data complexity by focusing on key

features. So you can use scikitle learn

to try out these algorithms on sample

data sets. So this will give you

hands-on experience with model training

and you will learn to fine-tune them to

get better results. Now before moving

on, if you are interested in advancing

your career in the field of AI and

machine learning, simple learns

post-graduate program delivered in

collaboration with Purdue University and

IBM is a perfect opportunity. This

highly ranked program offers a

comprehensive curriculum covering

essential topics like machine learning,

deep learning, NLP, computer vision,

reinforcement learning, generative AI,

prompt engineering, and many more. With

hands-on experience to 25 plus projects

and access to 20 plus cutting edge

tools, you will gain the skills needed

to excel in today's competitive job

market. So join now and elevate your

expertise with the backing of Produce

academic excellence and IBM's

industry-leading insight. You can find

the course link in the description box

and pin comments. Now moving on to the

seventh month, you'll be building and

training models with advanced libraries.

So by now you have experimented with

some basic models. So let's step it up

with advanced tools like TensorFlow and

PyTorch. So these libraries offer more

flexibility and power. So TensorFlow and

PyTorch. So here you can start with

simple models and work your way up. So

these libraries allow for building

neural networks which you'll be studying

more on the next month. Now once you

have become familiar with TensorFlow and

PyTorch, you can move on to model

training and evaluation. So you have to

learn to split data into training and

testing sets and evaluate models using

metrics like accuracy and precision. So

your goal this month should be to get

comfortable with these libraries and

understand how they handle data and

model training behind the scenes. So

once you are done with this, you'll be

moving on to the eighth month where

you'll be dealing with advanced machine

learning. So this month's concept will

be number one on n symbol learning which

means combining multiple models to get

better predictions. So here you'll be

learning about bagging for example

random forests here multiple decision

trees make predictions and then you have

boosting like ada boost xg boost so

models learn from each other's mistakes

over here and after ensemble learning

comes deep learning. So here you explore

neural networks which mimic the human

brain. So you'll learn about neural

network basics. So you can start with

simple fully connected networks and then

you can move on to back propagation and

gradient descent. So these helps your

model learn and improve. So you can use

TensorFlow or PyTorch to practice

building neural networks. So you can

work on projects to reinforce these

concepts. Now moving on, you have two

specialize on topics like NLP and

computer vision. So machine learning

applications are so powerful and here

you'll get a taste of two major fields

which is NLP or natural language

processing. So here they work with text

data with tasks like sentiment analysis

and text classification. So you can

start with basic pre-processing like

tokenization, stop word removal and move

to building simple NLP models. After

that you can try computer vision. So for

image data you have to learn CNN

convolutional neural networks. So these

network analyze visual patterns making

them ideal for image classification. So

you practice with open data sets like

text, documents or images and apply the

concepts you will learn to see results

in real world applications. Now in the

10th month you'll be dealing with model

deployment which is bringing your models

to life. So here you'll be using Flask

or Django. So you can use these

frameworks to create a web API so users

can interact with your model. For

example, build a web app that lets

people upload images for classification.

And then you can also try out Docker. So

package your model and its dependencies

so it can run on any machine. So this is

super helpful for deploying models

without compatibility issues. So by the

end of this month, you'll be able to

share your models with the world. So

moving on to the 11th month, you'll be

starting with cloud and production. So

this month, you'll learn how to deploy

models on the cloud and ensure they

perform well in real world environments.

So you'll be dealing with cloud

platforms like AWS, Google Cloud or

Azure. So you have to learn to deploy

models of the cloud provider

accessibility and scalability. And then

comes monitoring and maintenance. So

understand how to track your models

performance over time and update it as

needed. So these skills are essential

for maintaining models in production and

ensuring they stay reliable. And finally

you will be creating real world projects

and portfolio building. So here you have

to choose topics that interest you and

showcase your skills. So first you can

start with full projects. So complete

projects that go from data cleaning and

model building to deployment. So ideas

could be a sentiment analysis tool or an

image recognition app. And then you have

to build your portfolio. So organize and

document your projects, host them on

GitHub and create an online portfolio to

share with potential employers or

collaborators. So by following this road

map, you'll be well prepared to handle

real world machine learning challenges

and have an impressive portfolio to show

for it.

>> Welcome to machine learning tutorial

part two. My name is Richard Kersner

with the SimplyLearn team. That is

www.simplearn.com.

Get certified, get ahead. Today in our

second tutorial, we're going to cover K

means linear regression along with going

over the quiz questions we had during

our first tutorial. What's in it for

you? We're going to cover clustering.

What is clustering? K means clustering

which is one of the most common used

clustering tools out there including a

flowchart to understand K means

clustering and how it functions and then

we'll do an actual Python live demo on

clustering of cars based on brands. Then

we're going to cover logistic

regression. What is logistic regression?

Logistic regression curve and sigmoid

function. And then we'll do another

Python code demo to classify a tumor as

malignant or benign based on features.

And let's start with clustering. Suppose

we have a pile of books of different

genres. Now we divide them into

different groups like fiction, horror,

education, and as we can see from this

young lady, she definitely is into heavy

horror. You can just tell by those eyes

and the maple Canadian leaf on her

shirt. But we have fiction, horror, and

education. And we want to go ahead and

divide our books up. Well, organizing

objects into groups based on similarity

is clustering. And in this case, as

we're looking at the books, we're

talking about clustering things with

known categories. But you can also use

it to explore data. So you might not

know the categories. You just know that

you need to divide it up in some way to

conquer the data and to organize it

better. But in this case, we're going to

be looking at clustering in specific

categories. And let's just take a deeper

look at that. We're going to use K means

clustering. K means clustering is

probably the most commonly used

clustering tool in the machine learning

library. K means clustering is an

example of unsupervised learning. If you

remember from our previous thing, it is

used when you have unlabeled data. So we

don't know the answer yet. We have a

bunch of data that we want to cluster to

different groups. Define clusters in the

data based on feature similarity. So

we've introduced a couple terms here.

We've already talked about unsupervised

learning and unlabeled data. So we don't

know the answer yet. We're just going to

group stuff together and see if we can

find an unanswer

connect. We've also introduced feature

similarity. Features being different

features of the data. Now, with books,

we can easily see fiction and horror and

history books. But a lot of times with

data, some of that information isn't so

easy to see right when we first look at

it. And so, K means is one of those

tools where we can start finding things

that connect that match with each other.

Suppose we have these data points and

want to assign them into a cluster. Now

when I look at these data points, I

would probably group them into two

clusters just by looking at them. I'd

say two of these group of data kind of

come together. But in K means we pick K

clusters and assign random centrids to

clusters where the K clusters represents

two different clusters. We pick K

clusters and say random centroidids to

the clusters. Then we compute distance

from objects to the centrids. Now we

form new clusters based on minimum

distances and calculate the centrids. So

we figure out what the best distance is

for the centrid. Then we move the

centrid and recalculate those distances.

Repeat previous two steps iteratively

till the cluster centroid stop changing

their positions and become static.

Repeat previous two steps iteratively

till the cluster centroid stop changing

and the positions become static. Once

the clusters become static, then K means

clustering algorithm is said to be

converged. And there's another term we

see throughout machine learning is

converged. That means whatever math

we're using to figure out the answer has

come to a solution or it's converged on

an answer. Shall we see the flowchart to

understand make a little bit more sense

by putting it into a nice easy step by

step? So we start, we choose K. We'll

look at the elbow method in just a

moment. We assign random centrids to

clusters and sometimes you pick the

centrids because you might look at the

data in a in a graph and say ah these

are probably the central points. Then we

compute the distance from the objects to

the centrids. We take that and we form

new clusters based on minimum distance

and calculate their centrids. Then we

compute the distance from objects to the

new centrids. And then we go back and

repeat those last two steps. We

calculate the distances. So as we're

doing it, it brings into the new centrid

and then we move the centrid around and

we figure out what the best which

objects are closest to each centrid. So

the objects can switch from one centroid

to the other as the centroidids are

moved around and we continue that until

it is converged. Let's see an example of

this. Suppose we have this data set of

seven individuals and their score on two

topics A and B. Uh so here's our subject

in this case referring to the person

taking the uh test and then we have

subject A where we see what they've

scored on their first subject and we

have subject B and we can see what they

score on the second subject. Now let's

take two farthest apart points as

initial cluster centroidids. Now

remember we talked about selecting them

randomly or we can also just put them in

different points and pick the furthest

one apart so they move together. Either

one works okay depending on what kind of

data you're working on and what you know

about it. So we took the two furthest

points one and one and five and seven.

And now let's take the two farthest

apart points as initial cluster

centrids. Each point is then assigned to

the closest cluster with respect to the

distance from the centrids. So we take

each one of these points in there. We

measure that distance. And you can see

that if we measured each of those

distances and you use the the

Pythagorean theorem for a triangle in

this case because you know the x and the

y and you can figure out the diagonal

line from that or you can just take a

ruler and put it on your monitor. That'd

be kind of silly but it would work if

you're just eyeballing it. You can see

how they naturally come together in

certain areas. Now we again calculate

the centroidids of each cluster. So

cluster one and then cluster two and we

look at each individual dot. There's

one, two, three. We're in one cluster.

Uh the centrid then moves over. It

becomes 1.8 comma 2.3. So remember it

was at 1 and one. Well, the very center

of the data we're looking at would put

it at the one point roughly 22, but 1.8

and 2.3. And the second one, if we

wanted to make the overall mean vector,

the average vector of all the different

distances to that centrid, we come up

with 4, 1, and 54. So we've now moved

the centrids. We compare each

individual's distance to its own cluster

mean and to that of the opposite cluster

and we find build a nice chart on here

that the as we move that centrid around

we now have a new different kind of

clustering of groups and using uklidian

distance between the points and the mean

we get the same formula you see new

formulas coming up. So we have our

individual dots distance to the mean

centrid of the cluster and distance to

the mean centrid of the cluster. Only

individual three is nearer to the mean

of the opposite cluster cluster two than

its own cluster one. And you can see

here in the diagram where we've kind of

circled that one in the middle. So when

we've moved the clust the centroidids of

the clusters over one of the points

shifted to the other cluster because

it's closer to that group of

individuals. Thus, individual 3 is

relocated to cluster two, resulting in a

new partition. And we regenerate all

those numbers of how close they are to

the different clusters. For the new

clusters, we will find the actual

cluster centroidids. So now we move the

centrids over. And you can see that

we've now formed two very distinct

clusters on here. On comparing the

distance of each individual's distance

to its own cluster mean and to that of

the opposite cluster, we find that the

data points are stable. Hence, we have

our final clusters. Now if you remember

I brought up a concept earlier K mean on

the K means algorithm choosing the right

value of K will help in less number of

iterations and to find the appropriate

number of clusters in a data set we use

the elbow method and within sum of

squares WSS is defined as the sum of the

squared distance between each member of

the cluster and its centrid and so you

see we've done here is we have the

number of clusters and as you do the

same K means algorithm over the

different clusters and you calculate

what that centrid looks like and you

find the optimal you can actually find

the optimal number of clusters using the

elbow the graph is called as the elbow

method and on this we guessed at two

just by looking at the data but as you

can see the slope you actually just look

for right there where the elbow is in

the slope and you have a clear answer

that we want two different to start with

k means equals two a lot of times people

end up computing k means equals 2 3 four

five until they find the value which

fits on the elbow joint. Sometimes you

can just look at the data and if you're

really good with that specific domain

remember domain I mentioned that last

time you'll know that that where to pick

those numbers and where to start

guessing at what that k value is. So

let's take this and we're going to use a

use case using k means clustering to

cluster cars into brands using

parameters such as horsepower, cubic

inches, make, year, etc. So, we're going

to use the data set cars data having

information about three brands of cars,

Toyota, Honda, and Nissan. We'll go back

to my favorite tool, the Anaconda

Navigator with the Jupiter notebook. And

let's go ahead and flip over to our

Jupyter notebook. And in our Jupyter

Notebook, I'm going to go ahead and just

paste the uh basic code that we usually

start a lot of these off with. We're not

going to go too much into this code

because we've already discussed numpy.

We've already discussed mapplot library

and pandas. Numpy being the number

array, pandas being the pandas data

frame and mattplot for the graphing. And

don't forget uh since if you're using

the Jupyter notebook, you do need the

mattplot library in line so that it

plots everything on the screen. If

you're using a different Python editor,

then you probably don't need that

because it'll have a popup window on

your computer. And we'll go ahead and

run this just to load our libraries and

our setup into here. The next step is of

course to look at our data which I've

already opened up in a spreadsheet. And

you can see here we have the miles per

gallon, cylinders, cubic inches,

horsepower, weight pounds, how you know

how heavy it is, time it takes to get to

60. My card is probably on this one at

about 80 or 90. What year it is? So this

is you can actually see this is kind of

older cars and then the brand Toyota,

Honda, Nissan. So the different cars are

coming from all the way from 1971 if we

scroll down to uh the 80s. We have

between the 70s and 80s a number of cars

that they've put out. And let's uh we

come back here. We're going to do

importing the data. So we'll go ahead

and do data set equals and we'll use

pandas to read this in. And it's uh from

a CSV file. Remember, you can always

post this in the comments and request

the data files for these either in the

comments here on the YouTube video or go

to simplylearn.com and request that. The

car CSV, I put it in the same folder as

the code that I've stored. So, my Python

code is stored in the same folder, so I

don't have to put the full path. If you

store them in different folders, you do

have to change this and double check

your name variables. And we'll go ahead

and run this. And uh we've chosen data

set arbitrarily because, you know, it's

a data set we're importing. And we've

now imported our car CSV into the data

set. As you know, you have to prep the

data. So, we're going to create the X

data. This is the one that we're going

to try to figure out what's going on

with. And then there is a number of ways

to do this, but we'll do it in a simple

loop so you can actually see what's

going on. So, we'll do for i and x.c

columns. So, we're going to go through

each of the columns. And a lot of times

it's important I I'll make lists of the

columns and do this because I might

remove certain columns or there might be

columns that I want to be processed

differently. But for this we can go

ahead and take x of i and we want to go

fill na and that's a pandas command. But

the question is what are we going to

fill the missing data with? We

definitely don't want to just put in a

number that doesn't actually mean

something. And so one of the tricks you

can do with this is we can take x of i.

And in addition to that, we want to go

ahead and turn this into an integer

because a lot of these are integers. So

we'll go ahead and keep it integers. And

me add the bracket here. And a lot of

editors will do this. They'll think that

you're closing one bracket. Make sure

you get that second bracket in there if

it's a double bracket. That's always

something that happens regularly. So

once we have our integer of x of yi,

this is going to fill in any missing

data with the average. And I was so busy

closing one set of brackets, I forgot

that the mean is also has brackets in

there for the pandas. So we can see

here, we're going to fill in all the

data with the average value for that

column. So if there's missing data is in

the average of the data it does have.

Then once we've done that, we'll go

ahead and loop through it again

and just check and see to make sure

everything is filled in correctly. And

we'll print and then we take x is null.

And this returns a set of the null value

or the how many lines are null. And

we'll just sum that up to see what that

looks like. And so when I run this and

so with the X, what we want to do is we

want to remove the last column because

that had the models. That's what we're

trying to see if we can cluster these

things and figure out the models. There

is so many different ways to sort the X

out. For one, we could take the X and we

could go data set, our variable we're

using, and use the eyelocation, one of

the features that's in pandas, and we

could take that and then take all the

rows and all but the last column of the

data set. And at this time, we could do

values. We just convert it to values.

So, that's one way to do this. And if I

let me just put this down here and print

X, it's a capital X we chose. and I run

this, you can see it's just the values.

We could also take out the values and

it's not going to return anything

because there's no values connected to

it. What I like to do with this is

instead of doing the location which does

integers more common is to come in here

and we have our data set and we're going

to do data set dot or data set columns.

And remember that lists all the columns.

So if I come in here, let me just mark

that as red and I print data set.c

columns.

You can see that I have my index here. I

have my MPG cylinders everything

including the brand which we don't want.

So the way to get rid of the brand would

be to do data columns of everything but

the last one minus one. So now if I

print this, you'll see the brand

disappears. And so I can actually just

take data set columns minus one and I'll

put it right in here for the columns

we're going to look at.

And let's unmark this.

And unmark this.

And now if I do an x.ad

I now have a new data frame. And you can

see right here we have all the different

columns except for the brand at the end

of the year. And it turns out when you

start playing with the data set, you're

going to get an error later on and it'll

say cannot convert string to float

value. And that's because it for some

reason these things the way they

recorded them must have been recorded as

strings. So we have a neat feature in

here on pandas to convert. And it is

simply convert objects.

And for this we're going to do convert

oops convert underscore

numeric numeric equals true. And yes, I

did have to go look that up. I don't

have it memorized the convert numeric in

there. If I'm working with a lot of

these things, I remember them, but um

depending on where I'm at, what I'm

doing, I usually have to look it up. And

we run that. Oops, I must have missed

something in here. Let me double check

my spelling. And when I double check my

spilling, you'll see I missed the first

underscore in the convert objects. And

when I run this, it now has everything

converted into a numeric value because

that's what we're going to be working

with is numeric values down here.

And the next part is that we need to go

through the data and eliminate null

values. Most people when they're doing

small amounts, you working with small

data pools discover afterwards that they

have a null value and they have to go

back and do this. So, you know, be aware

whenever we're formatting this data,

things are going to pop up and sometimes

you go backwards to fix it. And that's

fine. That's just part of exploring the

data and understanding what you have.

And I should have done this earlier, but

let me go ahead and increase the size of

my window one notch.

There we go. Easier to see.

So, we'll do 4 I in working with X dot

columns. will page through all the

columns. And we want to take X of I and

we're going to change that. We're going

to alter it. And so with this, we want

to go ahead and fill in X of I. Pandas

has the fill in a. And that just fills

in any non-existent missing data. And

we'll put my brackets up. And there's a

lot of different ways to fill this data.

If you have a really large data set,

some people just void out that data

because if and then look at it later in

a separate exploration of data. One of

the tricks we can do is we can take our

column and we can find the means

and the means is in there or quotation

marks. So we take the columns, we're

going to fill in the non-existing one

with the means. The problem is that

returns a decimal float. So some of

these aren't decimals. Certainly, you

may need to be a little careful of doing

this, but for this example, we're just

going to fill it in with the integer

version of this. Keeps it on par with

the other data that isn't a decimal

point.

And then what we also want to do is we

want to double check. A lot of times you

do this first part first to double

check, then you do the fill, and then

you do it again just to make sure you

did it right. So, we're going to go

through and test for missing data. And

one of the re ways you can do that is

simply go in here and take our X of I

column. So it's going to go through the

X of I column. It says is null. So it's

going to return any any place there's a

null value. It actually goes through all

the rows of each column is null. And

then we want to go ahead and sum that.

So we take that, we add the sum value.

And these are all pandas. So is null is

a panda command and so is sum. And if we

go through that and we go ahead and run

it

and we go ahead and take and run that,

you'll see that all the columns have

zero null values. So we've now tested

and double checked and our data is nice

and clean. We have no null values.

Everything is now a number value. We

turned it into numeric and we've removed

the last column in our data. And at this

point, we're actually going to start

using the elbow method to find the

optimal number of clusters. So, we're

now actually getting into the sklearn

part. Uh, the K means clustering on

here. I guess we'll go ahead and zoom it

up one more notch so you can see what

I'm typing in here.

And then from sklearn going to or

sklearn

cluster, we're going to import K means.

I always forget to capitalize the K and

the M when I do this. So it's capital K,

capital M K means.

And we'll go and create a um array WCSS

equals we'll make it an empty array. If

you remember from the elbow method from

our slide

within the sums of squares, WSS is

defined as the sum of squared distance

between each member of the cluster and

it centrid. So we're looking at that

change in differences as far as a

squared distance. And we're going to run

this over a number of K mean values.

In fact, let's go for I in range. We'll

do 11 of them.

Range zero of 11.

And the first thing we're going to do is

we're going to create the actual we'll

do it all lowercase.

And so we're going to create this object

from the K means that we just imported.

And the variable that we want to put

into this is in clusters. We're going to

set that equals to I. That's the most

important one because we're looking at

how increasing the number of clusters

changes our answer. There are a lot of

settings to the K means. Our guys in the

back did a great job just kind of

playing with some of them. The most

common ones that you see in a lot of

stuff is how you enit your K means. So

we have K means plus plus. This is just

a tool to let the model itself be smart

how it picks it centrids to start with

its initial centroidids. We only want to

iterate no more than 300 times. We have

a max iteration we put in there. We have

the infinite the random state equals

zero. You really don't need to worry too

much about these when you're first

learning this. As you start digging in

deeper, you start finding that these are

shortcuts that will speed up the process

as far as a setup. But the big one that

we're working with is the inclusters

equals I. So, we're going to literally

train our K means 11 times. We're going

to do this process 11 times. And if

you're working with big data, you know,

the first thing you do is you run a

small sample of the data so you can test

all your stuff on it. And you can

already see the problem that if I'm

going to iterate through a terabyte of

data 11 times and then the K means

itself is iterating through the data

multiple times. That's a heck of a

process. So you got to be a little

careful with this. A lot of times though

you can find your elbow using the elbow

method. Find your optimal number on a

sample of data especially if you're

working with larger data sources. So we

want to go ahead and take our K means

and we're just going to fit it. If

you're looking at any of the sklearn,

very common that you fit your model. And

if you remember correctly, our variable

we're using is the capital X. And once

we fit this value, we go back to the um

array we made. And we want to go and

just append that value on the end.

And it's not the actual fit we're

pinning in there. It's when it generates

it, it generates the value you're

looking for is inertia. So k

means.inertia will pull that specific

value out that we need.

And let's get a visual on this. We'll do

our PLT plot. And what we're plotting

here

is first the x axis, which is range 0

11. So that will generate a nice little

plot there. And the wcss for our y axis.

It's always nice to give our uh plot a

title.

And let's see, we'll just give it the

elbow method for the title. And let's

get some labels. So let's go ahead and

do PLT X label.

And what we'll do, we'll do number of

clusters for that. And PLT Y label. And

for that, we can do oops, there we go.

WCSS since that's what we're doing on

the plot on there. And finally, we want

to go ahead and display our graph, which

is simply plt. Oops.

Show. There we go. And because we have

it set to inline, it'll appear inline.

Hopefully I didn't make a type error on

there.

And you can see we get a very nice

graph. You can see a very nice elbow

joint there at uh two and again right

around three and four. And then after

that there's not very much. Now as a

data scientist, if I was looking at

this, I would do either three or four.

And I'd actually try both of them to see

what the u output look like. And they've

already tried this in the back. So,

we're just going to use three as a setup

on here. And let's go ahead and see what

that looks like when we actually use

this to show the different kinds of

cars.

And so, let's go ahead and apply the K

means to the cars data set. And

basically, we're going to copy the code

that we loop through up above where K

means equals K means number of clusters.

And we're just going to set the number

of clusters to three since that's what

we're going to look for. And you could

do three and four on this and graph them

just to see how they come up

differently. It'd be kind of curious to

look at that. But for this, we're just

going to set it to three. Go ahead and

create our own variable Y k means for

our answers. And we're going to set that

equal to Whoops, my double equal there

to K means. But we're not going to do a

fit. We're going to do a fit predict is

the setup you want to use. And when

you're using untrained models, you'll

see um a slightly different because

usually you see fit and then you see

just the predict. But we want to both

fit and predict the k means on this. And

that's fit underscore predict. And then

our capital x is the data we're working

with.

And before we plot this data, we're

going to do a little pandas trick. We're

going to take our x value and we're

going to set x as matrix. So we're

converting this into a nice rows and

columns kind of setup. But we want the

we're going to have columns equals none.

So it's just going to be a matrix of

data in here. And let's go ahead and run

that.

A little warning. You'll see this

warnings pop up because things are

always being updated. So there's like

minor changes in the versions and future

versions. Let's set a matrix. Now that

it's more common to set it values

instead of doing as matrix, but mass

matrix works just fine for right now and

you'll want to update that later on. But

let's go ahead and dive in and plot this

and see what that looks like. And before

we dive into plotting this data, I

always like to take a look and see what

I am plotting. So let's take a look at

why K means. I'm just going to print

that out down here. And we see we have

an array of answers. We have 2 1 0 2 1

2. So it's clustering these different

rows of data based on the three

different spaces it thinks it's going to

be.

And then let's go ahead and print X and

see what we have for X. And we'll see

that X is an array. It's a matrix. So we

have our different values in the array.

And what we're going to do, it's very

hard to plot all the different values in

the array. So we're only going to be

looking at the first two or positions

zero and one. And if you were doing a

full presentation in front of the board

meeting, you might actually do a little

different and and dig a little deeper

into the different aspects because this

is all the different columns we looked

at. But we'll only look at columns one

and two for this to make it easy. So

let's go ahead and clear this data out

of here and let's bring up our plot. And

we're going to do a scatter plot here.

So pl scatter.

And

this looks a little complicated. So

let's explain what's going on with this.

We're going to take the x values

and we're only interested in y of k

means equals 0, the first cluster. Okay?

And then we're going to take value zero

for the x-axis. And then we're going to

do the same thing here. We're only

interested in k means equals 0, but

we're going to take the second column.

So we're only looking at the first two

columns in our answer or in the data.

And then the guys in the back played

with this a little bit to make it

pretty.

And they discovered that it looks good

with a size equals 100. That's the size

of the dots. We're going to use red for

this one. And when they were looking at

the data and what came out, it was

definitely the Toyota on this. We're

just going to go ahead and label it

Toyota. Again, that's something you

really have to explore in here as far as

playing with those numbers and see what

looks good. We'll go ahead and hit enter

in there. And I'm just going to paste in

the next two lines, which is the next

two cars. And this is our Nissa and

Honda. And you'll see with our scatter

plot, we're now looking at where Y_K

means equals 1. And we want the zero

column and YK means equals 2. Again,

we're looking at just the first two

columns, zero and one. And each of these

rows then corresponds to Nissan and

Honda.

And I'll go ahead and hit enter on

there. And uh finally, let's take a look

and put the centrids on there. Again,

we're going to do a scatter plot.

And on the centrids, you can just pull

that from our K means, the uh model we

created cluster centers. And we're going

to just do um

all of them in the first number and all

of them in the second number, which is

01 because you always start with zero

and one.

And then they were playing with the size

and everything to make it look good.

We'll do a size of 300. We're going to

make the color yellow. And we'll label

them. It's always good to have some good

labels. Centroidids.

And then we do want to do a title. PLT

title.

And pop up there. PLT title. So you

always make want to make your graphs

look pretty. And we'll call it clusters

of car make. And one of the features of

the plot library is you can add a

legend. It'll automatically bring in it

since we've already labeled the

different aspects of the legend with

Toyota, Nissan, and Honda.

And finally, we want to go ahead and

show so we can actually see it. And

remember, it's in line. Uh so if you're

using a different editor that's not the

Jupyter notebook, you'll get a popup of

this. And you should have a nice set of

clusters here. So we can look at this

and we have a clusters of Honda in

green, Toyota in red, Nissan in purple.

And you can see where they put the

centroidids to separate them.

Now when we're looking at this, we can

also plot a lot of other different data

on here as far because we only looked at

the first two columns. This is just

column one and two or 01 as as you label

them in computer scripting. But you can

see here we have a nice clusters of car

making. and we were able to pull out the

data and you can see how just these two

columns form very distinct clusters of

data. So if you were exploring new data

you might take a look and say well what

makes these different almost going in

reverse you start looking at the data

and pulling apart the columns to find

out why is the first group set up the

way it is. Maybe you're doing loans and

you want to go, well, why is this group

not defaulting on their loans and why is

the last group defaulting on their

loans? And why is the middle group 50%

defaulting on their bank loans? And you

start finding ways to manipulate the

data and pull out the answers you want.

So now that you've seen how to use K

mean for clustering, let's move on to

the next topic. Now let's look into

logistic regression. The logistic

regression algorithm is the simplest

classification algorithm used for binary

or multiclassification problems. And we

can see we have our little girl from

Canada who's into horror books is back.

That's actually really scary when you

think about that with those big eyes. In

the previous tutorial, we learned about

linear regression, dependent and

independent variables. So to brush up,

y= mx + c. Very basic algebraic function

of uh y and x. The dependent variable is

the target class variable we are going

to predict. The independent variables X1

all the way up to XN are the features or

attributes we're going to use to predict

the target class. We know what a linear

regression looks like. But using the

graph, we cannot divide the outcome into

categories. It's really hard to

categorize 1.5, 3.6, 9.8. Uh for

example, a linear regression graph can

tell us that with increase in number of

hours studied, the marks of a student

will increase, but it will not tell us

whether the student will pass or not. In

such cases where we need the output as

categorical value, we will use logistic

regression. And for that, we're going to

use the sigmoid function. So you can see

here we have our marks 0 to 100, number

of hours studied. That's going to be

what they're comparing it to in this

example. And we usually form a line that

says y = mx + c. And when we use the

sigmoid function, we have p = 1 / 1 + e

the minus y, it generates a sigmoid

curve. And so you can see right here

when you take the ln, which is the

natural logarithm. I always thought it

should be nl, not ln. That's just the

inverse of uh e your e to the minus y.

And so we do this, we get ln of p 1 - p

= m * x + c. That's the sigmoid curve

function we're looking for. And we can

zoom in on the function and you'll see

that the function as it deres goes to

one or to zero depending on what your x

value is. And the probability if it's

greater than 0.5, the value is

automatically rounded off to one

indicating that the student will pass.

So if they're doing a certain amount of

studying, they will probably pass. Then

you have a threshold value at the 0.5.

It automatically puts that right in the

middle usually. And your probability if

it's less than 0.5, the value run it off

to zero indicating the student will

fail. So if they're not studying very

hard, they're probably going to fail.

This, of course, is ignoring the

outliers of that one student who's just

a natural genius and doesn't need any

studying to memorize everything. That's

not me, unfortunately. Have to study

hard to learn new stuff. problem

statement to classify whether a tumor is

malignant or B9. And this is actually

one of my favorite data sets to play

with because it has so many features and

when you look at them, you really are

hard to understand. You can't just look

at them and know the answer. So it gives

you a chance to kind of dive into what

data looks like when you aren't able to

understand the specific domain of the

data. But I also want you to remind you

that in the domain of medicine, if I

told you that my probability was really

good at classified things that say 90%

or 95% and I'm classifying whether

you're going to have a malignant or a B9

tumor, I'm guessing that you're going to

go get it tested anyways. So you got to

remember the domain we're working with.

So why would you want to do that if you

know you're just going to go get a

biopsy? Because you know it's that

serious. This is like an all or nothing.

just referencing the domain. It's

important. It might help the doctor know

where to look just by understanding what

kind of tumor it is. So it might help

them or aid them on something they

missed from before. So let's go ahead

and dive into the code and I'll come

back to the domain part of it in just a

minute. So use case and we're going to

do our normal imports here where we're

importing numpy, pandas, seabour, the

mattplot library and we're going to do

mattplot library in line since I'm going

to switch over to Anaconda. So, let's go

ahead and flip over there and get this

started. So, I've opened up a new window

in my Anaconda Jupyter Notebook. And by

the way, Jupyter Notebook, uh, you don't

have to use Anaconda for the Jupyter

Notebook. I just love the interface and

all the tools that Anaconda brings. So,

we got our import numpy aspy

number array. We have our pandas pd.

We're going to bring in Seabor to help

us with our graphs as SNS. So many

really nice tools in both Seabour and

Mattplot library. And we'll do our

mapplot library.pipplot as plt. And then

of course we want to let it know to do

it in line. And let's go and just run

that. So it's all set up. And we're just

going to call our data data. Not

creative today. Uh equals pd. And this

happens to be in a CSV file. So we'll

use a pdread_csv.

And I happen to name the file. renamed

it data forp2.csv.

You can of course um write in the

comments below the YouTube and request

for the data set itself or go to the

SimplyLearn website and we'll be happy

to supply that for you. And let's just

um open up the data before we go any

further and let's just see what it looks

like in a spreadsheet.

So when I pop it open in a local

spreadsheet, this is just a CSV file,

comma separated variables. We have an

ID. So I guess the U categorizes for

reference or what ID which test was

done. The diagnosis M for malignant, B

for B9. So there's two different options

on there. And that's what we're going to

try to predict is the M and B and test

it. And then we have like the radius

mean or average the texture average,

perimeter mean, area mean, smoothness. I

don't know about you, but unless you're

a doctor in the field, most of the

stuff, I mean, you can guess what

concave means just by the term concave,

but I really wouldn't know what that

means in the measurements they're

taking. So, they have all kinds of stuff

like how smooth it is, uh, the symmetry,

and these are all float values. You just

page through them real quick, and you'll

see there's, I believe, 36, if I

remember correctly, in this one.

So there's a lot of different values

they take and all these measurements

they take when they go in there and they

take a look at the different growth, the

tumorous growth. So back in our data and

I put this in the same folder as a code.

So I saved this code in that folder.

Obviously if you have it in a different

location, you want to put the full path

in there and we'll just do uh pandas

first five lines of data with the data

head. And we run that. We can see that

we have pretty much what we just looked

at. We have an ID. We have a diagnosis.

If we go all the way across, you'll see

all the different columns coming across

displayed nicely for our data.

And while we're exploring the data, our

uh Seabor, which we referenced as SNS,

makes it very easy to go in here and do

a joint plot. You'll notice the very

similar to because it is sitting on top

of the U plot library. So, the joint

plot does a lot of work for us. And

we're just going to look at the first

two columns that we're interested in,

the radius mean and the texture mean.

We'll just look at those two columns and

data equals data. So that tells it which

two columns we're plotting and that

we're going to use the data that we

pulled in. Let's just run that. And it

generates a really nice graph on here.

And there's all kinds of cool things on

this graph to look at. I mean, we have

the texture mean and the radius mean

obviously the axes. You can also see

and uh one of the cool things on here is

you can also see the histogram. They

show that for the radius mean where is

the most common radius mean come up and

where the most common texture is. So

we're looking at the tech the on each

growth it's average texture and on each

radius it's average uh radius on there

gets a little confusing because we're

talking about the individual objects

average. And then we can also look over

here and see the the histogram showing

us the median or how common each

measurement is. And that's only two

columns. So let's dig a little deeper

into Seabor. They also have a heat map.

And if you're not familiar with heat

maps, a heat map just means it's in

color. That's all that means. Heat map.

I guess the original ones were plotting

heat density on something. And so ever

since then it's just called a heat map.

And we're going to take our data and get

our corresponding numbers to put that

into the heat map. And that's simply

data.coR

for that. That's a pandas expression.

Let's remember we're working in a pandas

data frame. So that's one of the cool

tools in pandas for our data. And let's

just pull that information into a heat

map and see what that looks like. And

you'll see that we're now looking at all

the different features. We have our ID.

We have our texture. We have our area,

our compactness, concave points. And if

you look down the middle of this chart

diagonal going from the upper left to

bottom right, it's all white. That's

because when you compare texture to

texture, they're identical. So they're

100% or in this case perfect one in

their correspondence.

And you'll see that when you look at say

area or right below it, it has almost a

black on there. when you compare it to

texture. So these have almost no

corresponding data. They don't really

form a linear graph or something that

you can look at and say how connected

they are. They're very scattered data.

This is really just a really nice graph

to get a quick look at your data.

Doesn't so much change what you do, but

it changes verifying. So when you get an

answer or something like that or you

start looking at some of these

individual pieces, you might go, "Hey,

that doesn't match. according to showing

our heat map, this should not correlate

with each other. And if it is, you're

going to have to start asking, well,

why? What's going on? What else is

coming in there? But it does show some

really cool information on here. I mean,

we can see from the ID, there's no real

one feature that just says if you go

across the top line that lights up.

There's no one feature that says, hey,

if the area is a certain size, then it's

going to be B9 or malignant. It says

there's some that sort of add up and

that's a big hint in the data that we're

trying to ID this whether it's malignant

or B9. That's a big hint to us as data

scientists to go okay we can't solve

this with any one feature. It's going to

be something that includes all the

features or many of the different

features to come up with a solution for

it. And while we're exploring the data

let's explore one more area and let's

look at data isnull. We want to check

for null values in our data. If you

remember from earlier in this tutorial,

we did it a little differently where we

added stuff up and sum them up. You can

actually with pandas do it really

quickly. Data.isnull and summit. And

it's going to go across all the columns.

So when I run this,

you're going to see all the columns come

up with no null data.

So we've just just to rehash these last

few steps. We've done a lot of

exploration. We have looked at the first

two columns and seen how they plot with

the seabour with a joint plot which

shows both the histogram and the data

plotted on the XY coordinates. And

obviously you can do that more in detail

with different columns and see how they

plot together. And then we took and did

the Seabor heat map the SNS

heat mapap of the data. And you can see

right here where it did a nice job

showing us some bright spots where stuff

correlates with each other and forms a

very nice combination or points of

scattering points. And you can also see

areas that don't.

And then finally, we went ahead and

checked the data. Is the data null

value? Do we have any missing data in

there? Very important step because it'll

crash later on. If you forget to do this

step, it will remind you when you get

that nice error code that says null

values. Okay. So, not a big deal if you

miss it, but it it's no fun having to go

back when you're when you're in a huge

process and you've missed this step and

now you're 10 steps later and you got to

go remember where you were pulling the

data in.

So, we need to go ahead and pull out our

X and our Y. So, we just put that down

here and we'll set the X equal to. And

there's a lot of different options here.

Certainly we could do X equals all the

columns except for the first two because

if you remember the first two is the ID

and the diagnosis. So that certainly

would be an option. But what we're going

to do is we're actually going to focus

on the worst. The worst radius, the

worst texture, parameter area,

smoothness, compactness, and so on. One

of the reasons to start dividing your

data up when you're looking at this

information is sometimes the data will

be the same data coming in. So if I have

two measurements coming into my model,

it might overweigh them. It might

overpower the other measurements because

it's measuring it's basically taking

that information in twice. That's a

little bit past the scope of this

tutorial. I want you to take away from

this though is that we are dividing the

data up into pieces and our team in the

back went ahead and said hey let's just

look at the worst. So I'm going to

create a an array and you'll see this

array radius worst texture worst

perimeter worst. We've just taken the

worst of the worst and I'm just going to

put that in my X. So this X is still a

pandas data frame but it's just those

columns. And our Y, if you remember

correctly, is going to be Oops, hold on

one second. It's not X. is data. There

we go. So, x equals data and then it's a

list of the different columns, the worst

of the worst. And if we're going to take

that, then we have to have our answer

for our y for the stuff we know. And if

you remember correctly, we're just going

to be looking at

the diagnosis. That's all we care about

is what is it diagnosed? Is it B9 or

malignant? And since it's a single

column, we can just do diagnosis. Oh, I

forgot to put the brackets. There we go.

Okay. So, it's just diagnosis on there.

And we can also real quickly do like an

X do. If you want to see what that looks

like and Y head

and run this and you'll see um it only

does the last one. I forgot about that.

If you don't do print, you can see that

the the Y.D is just mm because the first

ones are all malignant. And if I run

this, the X do head is just the first

five values of radius worst, texture

worst, parameter worst, area worst, and

so on. I'll go ahead and take that out.

So, moving down to the next step, we've

built our two data sets, our answer and

then the features we want to look at.

In data science, it's very important to

test your model. So we do that by

splitting the data

and from sklearn model selection we're

going to import train test split. So

we're going to split it into two groups.

There are so many ways to do this. I

noticed in one of the more modern ways

they actually split it into three groups

and then you model each group and test

it against the other groups. So you have

all kinds and there's reasons for that

which is past the scope of this and for

this particular example isn't necessary

for this. We're just going to split it

into two groups. one to train our data

and one to test our data. And the

sklearn uh.mmodel selection we have

train tests split. You could write your

own quick code to do this where you just

randomly divide the data up into two

groups but they do it for us nicely

and we actually can almost we can

actually do it in one statement with

this where we're going to generate four

variables capital X train capital X

test. So we have our training data we're

going to use to fit the model and then

we need something to test it and then we

have our y train. So we're going to

train the answer and then we have our

test. So this is the stuff we want to

see how good it did on our model. And

we'll go ahead and take our train test

split that we just imported.

And we're going to do X and our Y, our

two different data that's going in for

our split. And then the guys in the back

came up and wanted us to go ahead and

use a test size equals.3.

That's test size. Random state. It's

always nice to kind of switch a random

state around, but not that important.

What this means is that the test size is

we're going to take 30% of the data and

we're going to put that into our test

variables, our Y test and our X test.

And we're going to do 70% into the X

train and the Y train. So, we're going

to use 70% of the data to train our

model and 30% to test it. Let's go ahead

and run that and load those up. So now

we have all our stuff split up and all

our data ready to go. And now we get to

the actual logistics part. We're

actually going to do our create our

model. So let's go ahead and bring that

in from sklearn. We're going to bring in

our linear model and we're going to

import logistic regression. That's the

actual model we're using. And let's

we'll call it log model.

Oops, there we go. Model. And let's just

set this equal to our logistic

regression that we just imported. So now

we have a variable log model set to that

class for us to use. And with most the

uh models in the sklearn, we just need

to go ahead and fix it. Fit do a fit on

there. And we use our x train that we

separated out with our y train. And

let's go ahead and run this. So once

we've run this, we'll have a model that

fits this data that 70% of our training

data.

Uh, and of course it prints this out

that tells us all the different

variables that you can set on there.

There's a lot of different choices you

can make, but for Word do, we're just

going to let all the defaults set. We

don't really need to mess with those on

this particular example. And there's

nothing in here that really stands out

as super important until you start

fine-tuning it. But for what we're

doing, the basics will work just fine.

And then let's we need to go ahead and

test out our model. Is it working? So

let's create a variable Y predict. And

this is going to be equal to our log

model. And we want to do a predict.

Again, very standard format for the

sklearn library is taking your model and

doing a predict on it. And we're going

to test y predict against the y test. So

we want to know what the model thinks

it's going to be. That's what our y

predict is. And with that, we want the

capital xx test. So we have our train

set and our test set. And now we're

going to do our y predict. And let's go

ahead and run that.

And if we uh print

y predict, let me go ahead and run that.

You'll see it comes up and it predents a

prints a nice array of uh B and M for B9

and malignant

for all the different test data we put

in there. So, it does pretty good. We're

not sure exactly how good it does, but

we can see that it actually works and is

functional. Was very easy to create.

You'll always discover with our data

science that as you explore this, you

spend a significant amount of time

prepping your data and making sure your

data coming in is good. Uh there's a

saying, good data in, good answers out.

Bad data in, bad answers out. That's

only half the thing. That's only half of

it. Selecting your models becomes the

next part as far as how good your models

are. and then of course fine-tuning it

depending on what model you're using. So

we come in here, we want to know how

good this came out. So we have our Y

predict here, log model.predict X test.

So for deciding how good our model is,

we're going to go from the

sklearn.metrics,

we're going to import classification

report. And that just reports how good

our model is doing. And then we're going

to feed it the model data. And let's

just print this out. and we'll take our

uh classification report

and we're going to put into there

our test our actual data. So this is

what we actually know is true and our

prediction what our model predicted for

that data on the test side. And let's

run that and see what that does.

So we pull that up. You'll see that we

have um a precision for B9 and malignant

B and M. And we have a precision of 93

and 91, a total of 92. So it's kind of

the average between these two of 92.

There's all kinds of different

information on here. Your F1 score,

your recall, your support coming through

on this. And for this, I'll go ahead and

just flip back to our slides that they

put together for describing it. And so

here we're going to look at the

precision using the classification

report. And you see this is the same

print out I had up above. Some of the

numbers might be different because it

does randomly pick out which data we're

using. So this model is able to predict

the type of tumor with 91% accuracy. So

we look back here that's you will see

where we have uh B9 and malignant. It

actually has 92 coming up here. We're

looking about a 92 91% precision. And

remember I reminded you about domain.

So, when we're talking about the domain

of a medical domain with a very

catastrophic outcome, you know, at 91 or

92% precision, you're still going to go

in there and have somebody do a biopsy

on it. Very different than if you're

investing money and there's a 92% chance

you're going to earn 10% and 8% chance

you're going to lose 8%, you're probably

going to bet the money because at that

odds, it's pretty good that you'll make

some money. And in the long run, you do

that enough, you definitely will make

money. And also with this domain, I've

actually seen them use this to identify

different forms of cancer. That's one of

the things that they're starting to use

these models for because then it helps a

doctor know what to investigate. So that

wraps up this section. We're finally

we're going to go in there and let's

discuss the answers to the quiz asked in

machine learning tutorial part one. Can

you tell what's happening in the

following cases? Grouping documents into

different categories based on the topic

and content of each document. This is an

example of clustering where K means

clustering can be used to group the

documents by topics using bag of words

approach. So if you gotten in there that

you're looking for clustering and

hopefully you had at least one or two

examples like K means that are used for

clustering different things then give

yourself a two thumbs up. B identifying

handwritten digits in images correctly.

This is an example of classification.

The traditional approach to solving this

would be to extract digit dependent

features like curvature of different

digits etc. and then use a classifier

like SVM to distinguish between images.

Again, if you got the fact that it's a

classification example, give yourself a

thumb up. And if you're able to go, hey,

let's use SVM or another model for this,

give yourself those two thumbs up on it.

C. Behavior of a website indicating that

the site is not working as designed.

This is an example of anomaly detection.

In this case, the algorithm learns what

is normal and what is not normal,

usually by observing the logs of the

website. Give yourself a thumbs up if

you got that one. And just for a bonus,

can you think of another example of

anomaly detection? One of the ones I use

it for in my own business is detecting

anomalies in stock markets. Stock

markets are very fickled and they behave

very erratic. So finding those erratic

areas and then finding ways to track

down why they're erratic. Was something

released in social media? Was something

released you can see where knowing where

that anomaly is can help you to figure

out what the answer is to it in another

area. D predicting salary of an

individual based on his or her years of

experience. This is an example of

regression. This problem can be

mathematically defined as a function

between independent years of experience

and dependent variables salary of an

individual. And if you guess that this

was a regression model, give yourself a

thumbs up. And if you were able to

remember that it was between independent

and dependent variables and that terms,

give yourself two thumbs up. Summary. So

to wrap it up, we went over what is K

means and we went through also the chart

of choosing your elbow method and

assigning a random centrid to the

clusters, computing the distance and

then going in there and figuring out

what the minimum centroidids is and

computing the distance and going through

that loop until it gets the perfect

centrid. And we looked into the elbow

method to choose K based on running our

clusters across a number of variables

and finding the best location for that.

We did a nice example of clustering cars

with K means even though we only looked

at the first two columns to make it

simple and easy to graph. You can easily

extrapolate that and look at all the

different columns and see how they all

fit together. And we looked at what is

logistic regression. We discussed the

sigmoid function. What is logistic

regression? And then we went into an

example of classifying tumors with

logistics. I hope you enjoyed part two

of machine learning. So in today's

session we will discuss what RNN model

is. Moving ahead we will see why should

we use RNN. After that we will see how

does RNN work recurrent neural network.

After covering these topics we will move

forward and see types of RNN recurrent

neural network and applications of RNN.

At the end we will do a hands-off lab

demo of sentiment analysis using RNN. So

before starting let us have a simple

question to brush our knowledge. So

question is what are the application of

RNN? Okay, NLP,

time series, image captioning and all of

the above. Please answer in the comment

section below and we will update the

correct answer in the pin comments or

you can pause this video, give it a

thought and answer in the comment

section. Before we move on to the

programming part, let's discuss what RNN

is and proceed further for the same. So

what is RNN? Recurrent neural network.

So RNN work on the principle of saving

output on a particular layer and feeding

this back to the input in order to

predict the output of the layer. This is

how can convert a feed neural network

into a recurrent neural network RN. The

node in different layers of neural

network are compressed to form a single

layer of recurrent neural network. A B

and C are the parameters of neural

network. Now that you understand what

RNN is, let's look at the way why RNN.

Okay. So why RNN? RNN were created

because there are few issues in the feed

forward neural network cannot handle the

sequential data considers only the

current input cannot memorize previous

input. Okay. So the solution of these

issues is RNN and RNN can handle

sequential data accepting the current

input data and previously received input

data. So RNN can memorize previous input

due to their internal memory. So moving

forward let's see how does RNN networks

work. Okay. So the input layer X takes

an input to the neural network and

process it and the passes it into the

middle layer. The middle layer edge can

consist of multiple hidden layers each

with its own activation function and

weight and biases. If you have a neural

network where the various parameters of

different hidden layers are not affected

by the previous layer that is the neural

network does not have the memory then

you can use RNN. So the RNN will

standardize the different activation

function and weights and biases so that

each hidden layer has the same

parameter. Then instead of creating

multiple hidden layers, it will create

one end loop over it as many time it has

required. So moving forward let's see

types of RNN. So there are four types of

RNN

one to one,

one to many, many to many and many to

one.

So let's see one to one RNN. So this

type of neural network is known as the

vanilla neural network. It is used for

general machine learning problem which

has a single input and a single output.

Now see

one to many RNN. This type of neural

network has a single input and multiple

outputs. An example of this is a image

captioning. Now let's see many to one

RNN. This RNN take a sequence of input

and generates a single output. Sentiment

analysis is a good example of this kind

of neural network where a given sentence

can be classified as expressing positive

or negative sentiment. And the last one

is many to many RNN. This RNN takes a

sequence of inputs and generates a

sequence of output. Machine translation

is the one of the example. So moving

forward, let's see application of

recurrent neural network. First one is

image captioning. RNNs are used to

caption an image by analyzing the

activities present. The second one is

time series prediction. Any time series

problem like predicting the prices of

stocks in a particular month can be

solved using RNN. And the third one is

natural language processing. Text mining

and sentiment analysis can be carried

out using RNN or NLP. Natural language

processing. The fourth one is machine

translation. Given an input in one

language, RNNs can be used to translate

the input into different language as

output. So now let's move to the

programming part. First we will import

some libraries major libraries for the

first we will import for the data frame.

So I will write import

pd.

The second one is import numpy

as np.

So pandas is a software library written

for the python programming language for

data manipulation and analysis. In

particular, it offers a data structure

and operations for manipulating

numerical tables and the time series.

And this numpy numpy is a library for

the Python programming language adding

support to four large multi-dimensional

array and matrices along with a large

collection of highle mathematical

function to operate on these arrays.

Okay. So for plotting we will import

some libraries like seabon

as

SNS. This is nothing just a short form

of we don't have to write again and

again CON c we can write SNS. So then

another one is from

wordcloud

port

mattplot lib

dot

pip plot

s plt. library.

Okay, so Seabone is a library that uses

Matt plot lib underneath to plot graphs.

It will be used to visualize zandom

distribution and the word cloud is a

visual representations

of words. Cloud creators are used to

highlight popular words and phrases

based on frequency and relevance. They

provide you with quick and simple visual

insights that can lead to more in-depth

analysis. And this mattplot lil mattplot

lib is a plotting library for the python

programming language and its numerical

mathematic

ext extension numpy. It provides an

object- oriented API for embedding plots

into application using general purpose

UI.

Okay. Like tinker wxython QT or gtk.

So let's import some

NLTK

natural language toolkit.

So

I will write import

NLTK.

Okay. from

NLTK

dot stem

importizer

then from analytic dot corpus

imports

and

from

NL ticket dot tokenize

port

tokenize

NLTK the natural language toolkit or

more commonly NLTK is a suit of

libraries and programs for symbolic and

statical natural language processing for

English written in Python programming

language and this is stop words. Stop

words are words that are so common they

are basically ignored by typical

tokenizers and this word tokenize is a

function in Python that splits a given

sentence into words using the analytical

library. Okay. So let's import some

scikitlearn

library. So for that I will write from

skarn

dot model

collection

import

train

test.

Okay. Then from skarn

dot feature

extraction

dot text import

vectorizer.

And then from

skarn dot matrices

matrix

import

confusion metric

classification.

Okay.

So, scikitlarn is a free source software

machine learning library for Python

programming language. It features

various classification, regression and

clustering algorithms including support

vector machine learning, logistic

regression and many others like random

forest classifier. And this train test

split method is used to split our data

into train and test set. First, we need

to divide our data into features like X

and Y labels. And this TF ID vectorzer

converts a collection of raw documents

into a matrix of TF features. The fast

text or what to vectorizer what

embedding Python implementation and this

confusion matrix. A confusion matrix is

a table that is used to define the

performance of a classification

algorithm. Okay.

Then we'll import some libraries like

prom skarn

do linear model

port

logistic

regression.

So then from

colonm

port

and from

import

random

forests classifier.

Okay. Then from

skarn dot name base

portoli

base.

Okay.

So everything is correct. You will see

while running. So logistic regression

estimate the probability of an event

occurring such as voted or didn't vote

based on a given data set of the

independent variable.

L SVC logistic regression estimate

sorry linear support vector machine SVC

is an algorithm that attempts to find a

hyper plane to maximize the distance

between classified samples and this

random forest classifier creates a set

of decision trees from a randomly

selected subset of the training set

and this Bernoli NBoli

name base is a part of the name base

family it is based on Bernoli

distribution ution and accept only

binary values that is zero or one.

So let's import some tensorflow. So

import

tensorflow

dot

compad dot v2 and

then import

tensorflow

data sets

as tfds.

So, TensorFlow is a free and open-source

library for machine learning and

artificial intelligence across a range

of task but has a particular focus on

training and inference of deep neural

networks. Okay, let's import warnings.

Nothing. Warning.

The warnings

import

string

import.

So everything is basic. Just let's see

the pickle. Typically is a Python is

primarily used in serializing and

deserializing a Python object structure.

Okay, let's run it. Let's see how many

error

after that we will load the data set and

uh we will go through data

visualization. Okay. Word cloud cannot

import name word cloud. Okay. C C will

be capital here.

random forest.

Okay, it's still loading here. Let's

see. Okay, so loading is done. So now

let's load the data set. So we'll write

data equals to PD dot

read

CSV

name

test.

So you can find this data set on the

description box below.

According to question

polarity

ID,

comma date,

comma

query

per you forget comma

polarity. Okay.

Seems fine.

Let me change this first

is using RNN.

Okay.

So here I will write data plus data dot

sample.

Let's

do one.

Okay. So, let me like brief uh tell you

that what we are going. Okay.

Let me brief you like what we will do in

this sentiment analysis using RNA. So in

this demo like you will see uh text

processing on Twitter data set and after

that we will perform different machine

learning algorithms on the data such as

logistic regression random forest

classifier SVC nas to classify positive

and negative dudes. After that I will

also build RNN recurrent neural network

which is the best fit for such textual

sentiment analysis. Okay. Since it's a

sequential data set which is requirement

for the RNN network. So let's dive into.

So now

we will see the data data visualization

data set details target like the

polarity of the tweets zero negative.

Okay. then the date like date of the

tweet and the polarity and the user that

what tweeted then the text okay so I

will write print

data set

data

shape

okay

let me first do like this. Yeah.

So there are 20

or you can say two like rows and six

number of columns. Okay. So it is a huge

data. I will you can find this data set

from the description box below. So here

let's see the data

and why I use head. Head is used for

like

for showing

top 10 rows of the data set. If you will

use tail instead of head, it will show

the last 10 rows of the data set. Okay.

Here polarity zero. Zero means negative

and four means positive. Okay. Like you

can consider 01.

This is ID, date, then query. then user

then the text.

Okay. So

here I will do data

clarity.

Okay. These are the 04. Okay.

Uniqueness. Zero means negative and the

four means positive. replacing the value

four as one for the ease of

understanding what I said to you you can

consider as 01. So data

polarity

to data

polarity

to one

and then data.

So now you can see 0 1 0 1 0 1 1 0.

Okay.

So if you will write only head it will

show the top five rows only. Okay.

So now let's use one Python function

describe

data dotribe.

So as you can see here count is two lakh

and the mean of the particular row is

this and the ID is this standard

deviation minimum value the 25% the 50%

and the 75% and the maximum

okay let's see the number of positive

versus negative tagged sentence okay

so here I I will write positives

to data

polarity

data dot polarity

= 1.

Then it is

data

polarity

data dot polarity

is equals to zero.

Print

total

length of the data is

dot format

data

dot shape. Yep.

Now I will print

the total length, the negative and the

positive. Okay. So number of positive

Okay.

Format

positives.

So I will copy

this and paste it here.

And here I will do the changes for the

negatives.

Okay. Now let's see.

So here polarity is not defined.

So as you can see the total length of

the data is two lakh and the number of

positive sentences is like one lakh 46

and number of negatives okay spelling

this

the number of negative text sentences

99,954.

Okay. So now we have a brief data.

So now let's get a word count p of text.

So for this I will write

count

words

done

length of

start split.

Okay.

And now let's plot a word count

distribution for both positive and

negative. So I will create a bar plot.

So for that I will write it

word

count.

data

text

dot apply

but count.

Okay, then I will write P positive= data

then

count

data dot polarity

is equals to 1

and

let me copy this Here

I will write zero

and okay then

plt dot figure

and figure size

equ= to

12 Thanks.

Okay. Then plt

LT dota

45

then plt dot x label

word count

plt dot y label

and frequency

we'll write uh g

dot

comma n

Uh

alpha also 0.5 Five

positive.

Okay. Then let's make a legend also.

Location should be

Right.

False.

Data word count equals to

Okay, my bad.

So as you can see the positive and the

negatives.

Okay.

So these are the like word count

distribution for both positive and

negative. Okay.

Now let's uh what we can do we can do

the get like get the common words in

training data set for the training data

set. So for that I will do

from

collections

import

counter

or

words

to

for

test

data

text

line

dot split

forward. Word

and words.

If length of

word

than two

all

dot

one dot lower

here I can write counter

all words

dot most

common then I need 20.

So as you can see these are the most

common word used like in every sentence

the and you for have that I am but just

like this out over all.

So these are the most common words like

it used the is used like 64,000 times

and like this UR is used for 8,000 times

something like that. So now we will do

some data pro data processing. Okay. Now

let's do the data processing.

So

div

and SNS dot current plot

data

polarity.

Okay,

these are the uh negatives and this

positives.

There is a slight change I guess that is

why it's not looking

so much of different like there's a

slight

46 different so that is why it's looking

almost same. Okay.

So now removing the unnecessary columns

like query, user, word count, data dot

drop,

date

query

and word count.

X = 1

comma

place= to true.

Okay.

Uh A will be true.

So here I will write data

what is this? No. Okay my bad.

So here I will write data dot drop

id

comma

one

then data dot head

the data see we have only the to the

polarity and the text. Okay.

So

now uh let's see the null values.

So data

dot

um

data

print. Okay.

So there is no null values. So now

converting pandas's object to a string

type.

For that we have to write

text

to data

text.

Yeah.

Get as type.

Yeah. So now download the stop words

NLTK.

Download

words.

words

as you said

stop words

it's in English

stop

This

These are some, you know, stop words.

So moving forward, let's download

NLTK dot download.net.

So the pre-processing steps taken are

like lower casting each text is

converted to lower case then remover of

URLs will do this we will do okay links

starting with http or https or ww are

replaced by like commas and removing

usernames removing short words removing

stop words like limitization is the

process will do of for the converting a

word to its base Okay. So for that

what I will do

we'll just copy the whole code for you.

We'll explain you one by one what I've

done.

Okay.

So this is a course for the URL pattern

for removing all the WW, HTTPS and HTTP

type of thing and removing

them. Then I have used pattern for the

lower casting removing all the URLs.

Okay.

Then removing all the usernames like at

the red and removing punctuations

and stop words.

Okay. Like this.

So now what we have to do data

processed

weights

then data

Next

dot apply

lambda

x

process

then tweets.

Okay.

Then print

next.

reprocessing.

It is taking time

It will be completed. It will return

here the text prep-processing is done.

Okay.

As you can see the text prep-processing

is done. So now let's check

data dot add

10.

As you can see see the at the rate and

this slices are gone.

Okay.

So now the text is pre-processed.

So now what we will do? We will analyze

the data. So now we are going to analyze

the pre-processed data to get an

understanding of it. We will plot word

clouds for positive and negative dudes

from our data set and see which words

occurs the most. Okay. First we will uh

create for the negative words or

negative tweets you can say. So I will

write pl dot figure

then figure size

15.

Okay. Then word cloud also

word cloud

x words

2,00 comma

width = to 1,600

comma

height = to 800

rate

dot join the data dot polarity

and I will write here polarity

okay equals equals to zero

again

then

processed tweets.

Okay.

Then here

I have to write plt dot show

me show

wcolation

linear.

Perhaps you forget the comma here.

So 2000

then comma width

dot generate

here. what I can do.

Let me run now. Let's see. Hope this

time it will work.

Guess is still loading.

As you can see this is

okay like today I am and work don't wish

they need much. These are the most

negative tweets. Okay, words from

negative tweets you can say,

right? So, let's see the positive

tweets. Okay,

so the thing will be same.

Let me copy

paste it here. So for this I will do one

it will take a little bit of time

to come loading like as you can see hit

can't. Okay sorry

these are the negative words.

Okay still loading. So let's wait for

like few seconds.

Now you can see the positive words like

love, okay, good, lol

and awesome something like that. Okay,

so these are some

positive words. So now let's do the

vectorzation and splitting the data like

storing into input variable process to X

and output variable polarity to Y. Okay,

we'll do that.

So x = to data

possessed

with

values

and pi= to data

entity

dot

values.

Okay.

Now I will write here print

dot shape

print y dot.shape.

Okay cool. So now what we will do we

will convert text to word frequency

vectors. Okay. TF to IDF. So this is an

acronym that stand for term frequency to

inverse document frequency which are the

components of the resulting scores

assigned to each word. Okay. So term

frequency this summarize how often a

given word appears within a document and

inverse document frequency this

downscales word that appear a lot across

documents. Okay. So now here we will

convert a collection of raw documents to

a matrix of TF to IDF features. Okay.

And then I will write

enter

kazut

riser

and sublinear

x = to

dot with

transform

printed.

Okay.

Print

number of feature

comma length

vector

do get

their names.

So number of feature words are like 1703

to 1.

Okay. Now we will do like

now let's print the shape.

So now we will do the split uh spread to

train and test. So the pre-provised data

is divided into two sets of data

training data and the testing data. So

data set upon which the model would be

trained on contains 80% data and the

test data is the data set upon which

model would be tested again contains 20%

of data. So for that I will write extra

test

comma

Test

test size

= to 0.20 2

random

state

101.

Okay. Random state.

So what I will do? I will do the you

know print the shape of X train, Y

train, X test, Y test like how many

columns are there? Rows not column

exactly the rows are there. Okay.

So we'll paste there. So see

extra train like this is a total was

like two lakh.

Okay. So 1 lakh 60,000 in training as we

discussed earlier like 80% in training

and 20% in testing. Okay.

So now let's do the model building.

Okay. Model evaluating functions. So now

let's make a model.

Okay.

And first I will do I will write and

then I will explain you the whole. Okay.

So here what I did uh this will tell you

the accuracy of the model of training

data and the testing data. Okay. Then we

will predict the values for test data

set and the evaluation for the data set.

Then we will compute and plot the

confusion matrix.

Okay, the both the categories negative

positives. Okay, group name will be true

negative and the false positive. Okay,

so there's nothing that's let's run it.

So now what we will do? We will do first

for the logistic regression. So here I

will write LG equals to

logistic

regression.

Okay. Then history

equals to LG do fit

X train,

Y train

with model

evaluate

LG. Now let's see

this is for the logistic regression.

Okay,

as you can see the accuracy of the

training data is 83% the testing data is

77%.

Okay.

So this is the confidence matrix the

predictive value like these are the

categories.

Now let's see for the linear SPM. For

that I will write SPM

equals to

SVC

then SVM

dot fit

train.

Then model

evaluate

of SVM.

Okay.

And after that we will do for random

forest and the N base. Okay. Then we

will start with the RNN.

So as you can see the accuracy of

training data is very pretty good 93%

and logation is 83% and the testing is

less than

regression model. Let's see for the

random forest. So I will write here RF

equals to

random forest

fire

m=

to 20

criterion = to

tropy Okay.

Then max

depth equals to 50.

Then RF dot fit

X train,

Y train

and model

evaluate.

Okay,

loading. Let's see the accuracy how it

will come.

After this we will do for the name base

and after that we will move on to the

our main model RNN recurrent neural

networks.

It's still loading.

Guess it will take little bit of time.

So as you can see the confusion matrix.

Okay. So training data accuracy is 75%

very less. So now let's see the last

model name base. Okay. So, NB equals to

NB

NB dot fit

SP,

wide train.

Okay. Then model

evaluate.

Oh, NAB base training 867.

So, as for

linear SEC has the best

test training uh accuracy you can say

and the best testing accuracy is 7670

76.45 4 five

see logistic regression. So now let's

move to the our main model RNN. So what

is RNN recurrent neural network at the

start

are the state-ofthe-art algorithm for

sequential data and are used by Apple CD

and Google search voice. It is the first

algorithm that remembers its input due

to an internal memory which make it

perfectly suited for machine learning

problem that involve sequential data.

And there is one more thing embedding

layer. Embedding layer is one of the

available layers in KAS. This is mainly

used in natural language processing

related applications such as language

modeling but it can also be used with

other tasks that involve neural networks

while dealing with NLP problems. We can

use pre-trained word embedding such as

glow.

Alternately we can also train our own

embeddings using kas emitting layer.

LSTM layer long short-term memory

networks usually called LSTMs I have

made already many videos you can check

it out were introduced by Skyer these

have widely been used for speech

recognition language processing

sentiment analysis and text prediction

before going deep into LSTM we should

first understand the need of LSTM which

can be explained by the drawback of

practical use of RNN so let's start with

RNA

Okay.

So here I will importing some libraries.

Okay. So after that I will write import

kas

version

2.110. Okay fine.

So now let's

paint

X test, comma,

white train.

Let's do train

test

weights, comma, data dot polarity

dot values.

Then test

size equals to 0.2. Test size 0.2 means

like 80 and 20%

thing 80 to training and then 20% to

testing.

Okay.

and let's

the model evaluation. Okay.

So I will these are relu sigmoid all the

you know the layers.

So now this epoch it will run till 5,000

like count will go till 5,000. Okay see

the 5,000 and it will go to 1 to 10. So

it will take time. So I will get back to

you after this completing this. Okay.

Now as you can see uh the box

ran successfully. Okay. So what should I

do? But I will give some space here. So

now we will see the positive and

negative outcome. Okay. This is

something like testing. Okay. We will

test. We will predict. we will give one

uh a sentence and then we will predict

it is coming right or wrong. The

accuracy is giving a right or wrong.

Okay. So here I will write

sequence

equals to tokenizer

dot text

to

sequences.

Okay, then I'll write this

data science

article.

This was

okay.

So here I will write test equals to P

sequences

and here I will write sequence

Comma max length

to

max length.

Then I will write here prediction equals

to model.

We write model

then we'll write model 12

dot predict

then test.

Okay.

If diction

is greater than 0.5 means 50%.

Then

it should print

positive.

Okay.

Else

negative.

Okay. Let me run this.

Okay. Sequential

object has no okay spellic

see the negative because here is the

word worst it is showing correct. Now

check from the RNN model. So model

equals to kas dot models dot load

models. Here we will load RNN model. RNN

model

SG file. It is pre-trained model. Okay.

Pretend RNN model. So sequence

tokenizer

dot text

to sequences.

Then

I will write here this this

ML

course

is best.

Okay. S equals to P sequences

sequence

X

okay then prediction equals to model dot

predict

Then test

if prediction is greater than 0.5

in

positive 0.5 means 50% more than 50%.

Else

print negative

attribute load models

positive because this ML course is best.

So there is no negative word.

Okay.

So what we will do now we will do model

saving loading and prediction. Okay. So

for that uh I will write import pickle

file = to open

vectorzer

then

here I will Pickle

dot dump

dump

vector file vector.

Okay.

So like this I have to write for name

base logitation SVM and random forest.

So

what I will do

right here.

Okay. Let's run this.

Okay.

Now what we have to do? We have to

predict using saved model. Okay.

What we will do here? We will load model

first and we will predict. Okay. So

first I will write the function name

load

models

and we will load the vectorzer. So file

equals to

open

vectorzer

dot pickle

IBizer

file.

file dot close.

Now I'm loading the logistic regression

model. So for that we have to write open

B

LG to pick

code

file

then file dot close

then

riser

LG.

Okay.

Yeah. So now we will predict the

sentiment. So for that I will write here

predict

riser

text.

Okay. Uh so here we will predict the

sentiment. So for that

text equals to

process

then demands for

sentiment in

text.

Then text

data

dot transform.

This is

okay. Then sentiment

model dot predict.

So here I will make a list of text with

sentiment. So for that I will write data

equals to empty array. Then for text

prediction

and zip

text x

sentiment

dt

append

text prediction.

Okay.

Then we will convert the list into pas

data frames. So for that I will write df

= to ad dot data plane

comma columns

person

next

comma

sentiment

then df equals to df dot

Replace

comma 1

positive

and here.

Okay.

So at last I will write here if

to

then here we will loading the model

vectorzer

comma ng plus load.

Here we text to classify like what

should be in the list. So like text

here I will like I love machine

name.

So

John

be so

So here df equals to date

the command text

then print

df.

I love machine learning. Positive. B is

so active. Positive. J I feel so good.

Negative. Okay. There is

and

yeah. See now it's coming. Okay.

This is how you can do the sentiment

analysis using uh RNN model. Here we

have loaded RNN model. So it is showing

right. So let's do them as poses

right here.

Add

one.

negative. Okay,

RNN model is working. So right, today we

are going to explore K nearest neighbors

or KN&N which is one of the most popular

algorithms in data science. Python is a

powerful tool for data science and KNN

is great for classifying data by

predicting the category of sample base

on its closest neighbor. This algorithms

is used in many field like healthcare,

finance and agriculture helping us make

decision based on data. The best part it

is really easy to use. You just need to

pick a number for K and choose a

distance function to compare data

points. However, KN&N has it downsides.

It doesn't work well with the large data

set and it require proper scaling of the

data to get accurate result. In this

video, we will show you how KNN work

with real data set, the Iris data set.

We'll walk you through simple Python

code and demonstrate how to find the

best K value to maximize your model's

accuracy. So stay tuned to seekn in

action. So welcome to the demo part. So

I'm here using Google Collab. So you can

use any of your favorite ID like Jupyter

notebook, Intelligi, Visual Code Studio,

anything. Okay. So let me rename this

file as KNN classification.

Okay, cool. So let me tell you that KN&N

can be used for the classification

regression predictive problems. So KN

falls in the supervised learning family

of algorithms. Okay, so we will measure

the distance between the K neighbors and

the first step will be we will choose

the number of K of neighbors. Then uh uh

we'll take the k nearest neighbors of

the new data point according to your

distance metric. And the step three will

be our among these case neighbors count

the number of data points of each

category. Okay. Step four we will assign

the new data points to the category

where you counted the most neighbors.

Okay. So let's start. First let's import

some library. import

numpy

as np and let's import

pandas sp. So everyone knows what is

numpy and the pandas. Okay, so numpy is

a library for the python programming

language adding support for uh you know

large multi- dimensional arrays. Okay.

Along with the large collection of uh

what to say uh highlevel mathematical

functions and various pandas is a

software library written for the Python

programming language for the data

manipulation and analysis uh all the

data frames and uh data structures it

offers for the manipulating numerical

tables. Okay. And the time series you

can see. So moving forward uh we'll

import our data set. So you can download

the data set from the description box

below. Okay. data set

equals to pb dot read

csv. The data name is iris dot csv.

Okay. This is how you read uh your data

in python. Okay. Yeah. Data set is

loaded. So data set dot shape

shape. Okay. Yeah. So data uh set dot

shape is used for how many numbers of

rows and columns present in your data

set. Okay, 150 rows and six columns. So

let me tell you brief about data set

this data set. So this data set include

three Iris species with uh 50 samples

each as well as some properties about

each flower. So one flower species is

linearly separable from the other two

you can say but the other two are not

linearly separable from each other.

Okay. And this shape I told you we can

get a quick idea of how many instances

of rows and columns are present in our

data set. So let's see our data set.

Data set dot

head. So head is used for uh you know by

you can see top five rows of your data

set using head and if you will use tail

you instead of head you can see the last

five rows of your data set. Cool. Okay.

So columns are ID sample length sample

width petal length petal width and the

species. Cool. Then moving forward let's

describe our data sets. So these are the

basics uh basic function. Okay.

of Python you can say.

So data set.escribed what describes do

is it will give you count of all the

rows mean value standard deviation value

minimum value what is the 25% okay of

all the values in the particular row.

What is the 50%? What is the 75%? What

is the maximum? Maximum is 150 you can

say. Okay. and 25% of 150 is 38.25 25

this okay of all the columns if it is

normal uh you know character so it won't

give you any data okay cool yeah so

moving forward uh let's now take a look

at the number of instances row belong to

each classes okay so we will write data

set dot

group by

species

dot size. Okay. Uh spec S is capital

that's why it's showing the error. Yeah.

So you can see Iris Satossa are 50, Iris

verical are 50 and virginica is 50.

Okay. And the data type type is integer.

Cool. So as you can see data set

contains six columns like ID, sample

length, sample width and petal length,

petal width and spacing. The actual

features are described by columns 1 to

four. the last columns labels or

samples. Okay. So firstly we need to

split data into two arrays like X

features and Y labels. So how we will do

this? By writing code like feature

columns equals to

sample length sample width petal length

petal width. Just remember you are

writing correct name. Okay. Then x = to

data set

feature

columns.

Okay. Dot

values.

Then y = to

data set

species dot values. Cool. Then let me

run it. So this is uh how we can split

the data set okay into two arrays X and

the Y. Okay. In X there are feature

columns. These four columns are there

and in Y species column is there. Okay.

And what is the species column this

Satossa venica and all this. Cool.

Then now we will do label encoding. So

as you can see labels are categorical

Kverse classifier does not accept string

labels. So we need to use label encoder

to transform them into numbers. Okay.

Then iris satossa correspond to zero.

Iris vericy color correspond to one and

I is virginica correspond to two. Okay.

012. Cool. So how we can write from

skarn dot pre-processing

import

label

encoder okay so what we'll do label

encoder transform them into numbers okay

so I will write

l equals to

label encoder

okay my bad label encoder. Okay. Then y

= to ele alate transform y. Okay. So now

I will run it. Yeah. Correct. So

splitting data set into training set and

the test set now. Okay. So now we'll

split data set into training set and

test set to check later on whether or

not a classifier work correctly or not.

Okay. So here I will write from skarn

dot not cross. I will write it here.

Skarn domodel selection

import

train test split. So what is train test

split? Train test split is a model

validation procedure that reveals how

your model performs on your new data.

Okay. And what is X train access Y train

Y test in Python. Okay. Let me first

write it and then I will let you know.

Okay. Then I will write here

x train comma x test comma y train comma

y test. Okay

train test

split

then x comma y

comma test size

0.2 and the random is this. Okay. Okay.

Some error came. Okay. Underscore model

selection.

Yeah. Cool. So what is X train X test? Y

train Y test. Okay. So X train and Y

train sets are used for training and

fitting the model. Okay. So the X test

and the Y test are the set used for

testing the model and it's predicting

the right outputs level. Okay. So here

you can see test size is 0.2. into means

80% is for testing or sorry 80% is for

training and 20% is for testing for the

new data. Cool. Yeah. So now we will see

some uh let's do some data

visualization. Okay. So here I will

write import

mattplot lib

dotpipplot

as plt

then import

cb

as sns

then here I will write person mattplot

lib

in line. So what is mattplot lib?

Mattplot lib is a plotting library for

the python programming language and it's

numerical mathematic extension. Okay. So

numpy it provides an object API for

embedding plots into application using

generating purpose GUI toolkits like

kintter and python or gtk. Okay. Whereas

seabon seabon is a library for making

statical graph in python. It builds on

the top of mattplot lip and integrates

closely with pandas data structure.

Okay, seborn helps us to explore and

understand the data. Okay, so I will run

it here. I will write from

pandas dot plotting

import

parallel

coordinates.

Then plt dot figure

size should be

15, 10. Okay.

Then parallel coordinates.

Okay. Then data set dot drop. I don't

need id,

x is one.

Okay. then comma spaces

then plt dot title

and let parallel

coordinates

plot okay and

you can give some font size

equals to 20

then font

weight

equals to

okay Let's add bold only bold then plt

dot x label

then

features

comma

font size

to 15 then plt

dot y label then I'll write here

features

values

comma font size

equals to 15. Okay. then plt dot legend

then locals to 1 comma

I will write have frame on

equals to true comma shadow

equals to true comma face color

equals to White

T should be capital

and comma edge color

equals to

okay plt dot show okay some error is

there plig

size okay spelling mistake Take.

Okay. One more error. PLT. Legend. Okay.

Face color. Okay. Some spelling stick.

So yeah, let me make it output in full

screen. Yeah. So parallel coordinates is

a plotting technique for plotting you

know multivariate data. So it allows one

to see clusters in the data and to

estimate other stat visually. So using

parallel coordinate points uh you know

are represented as the connected line

segments as you can see. Okay. And each

vertical line represent one attribute

and one set of connected line segments

represent one data point. Okay. And

points that tend to cluster will appear

closer together. Okay. So this uh this

color is iris satossa and this is irisy

color and this ba one is iris virginica

as you can see in the legend. Okay,

cool. So moving forward let's create

another graph and curves. Okay so here I

will write from

p and dotplotting import scatter. Oh,

this uh plotting import

Andreo curves.

Okay, then plt

dot figure. It's AI based. So it's

giving me suggestions. Suggestions.

Suggestions. Okay. So sometimes

suggestions are good but not always.

Yeah. Let's carry on. Figure size is 15,

10.

Then Andrew

curves

then I will add data set dot drop id

access spaces and plt and curves plot.

Okay fine. Okay, let me add we don't

need X label and all. Let me add legend

plot

legend.

Then same LOC equals to 1. Then

proposition size.

Then I will write here size

is 15

of frame on equals to two comma shadow

equals to true. So this is uh truly

based upon you if you want to add legend

or not or you can skip. If you want to

skip you can skip. Okay. Face color

equals to white. Then

edge color equals to black. Okay. Then

plt dot show.

Yeah. So no error is there. So let me

first view output in full screen. Okay.

So Andrew curves these are the endoc

curves. Okay. You can see the graph in

the curve. So and curve allow one to

plot multivariate data as a large number

of curves. So that are created using the

attributes of samples okay as

coefficient for 4year series. Okay. So

by coloring these curves differently for

each class. It is possible to visualize

data clustering. So curves belongs

belonging to samples of the same class

will usually be closer together and they

form large structure. Okay. As you can

see here and this is the legend why we

are setting the four color is should be

white and yeah and the edge color should

be black okay and frame on shadow should

be there you can see the shadow okay

like this if you want to skip you can

skip this part legend part legend part

okay but like it's good to have it's

good practice to have this cool so let's

create one small pair plot okay so I

will write here plt dot

figure

then I will write SNS dot pair plot

then I will write data set

dot drop

we'll drop again id we don't need comma

xis is one then I will write u equals to

species

size equals to three.

Then markers equals to

O SD. Yeah. Cool. Then plt dot show.

It's running. Yeah. So first let me make

it to full screen. Yeah. So here pair

wise is useful when you want to

visualize the distribution of the

variable or the relationship between the

multiple variable separately within

subsets of your data set. Okay. So here

this blue one is stoa verol and the

virginica. Okay. So these are some uh

graph you can say or here you can see

sample length. Okay. So green one is

virginica and versol are almost having

same saple length. Okay. And here are

the different types of graphs. Okay. If

you don't want to uh let's say if you

don't know how to read this graph you

can use this graph or you can use this

graph. Either you can use this graph.

Okay. That's the power of pair pair wise

plots you can say. And the sample width

cm. Okay. Satossa. No. Okay, virginica

and verticola again having almost same

sample width. Okay, and here as well.

Okay, this stoa is in different form.

Okay. Yeah. So, we'll uh create one

small uh pair of u one small graph then

we'll move forward. Okay. Uh I will

write here plt dot. So now we are

creating box plot figure.

Okay. Then data set dot drop.

Then again ID commais should be one

dot box plot. Okay. Then figure size I

chose this. Yeah. Let's run it. Yeah.

Again see this is the box plot. Petal

length same petal width sap length sle

width okay according to width it is

showing like from this to this width we

have the veric color and from this to

this width this length sorry sample

length we have uh that virginica lies

here Iris virginica and from here to

here iristosa lies okay the sample cool

and like you can use 3D models you can

use different types of charts. Okay, I

did four and four are enough to read the

data set. Okay, so now we will do uh

some cannon classification. Okay, we

will make prediction and we will see the

accuracy of our data set. Okay, so now

what I will do? I will write here from

skarn

dot

neighbors

import k neighbors. Okay.

Write from skarn dot

neighbors

import

k neighbor

classifier.

Okay.

Then I will write here from

skarn

dot matrix

import

confusion

matrix

comma accuracy

accuracy score. Okay, then I will write

from

skarn dot model selection

import

cross

value score. Okay, so here what we I did

uh we are fitting the classifier to the

training set and loading the libraries

basically. Okay, then we'll initiate a

learning model K equals to three. Okay.

So here I will write classifier.

Let me give one space. Classifier equals

to

neighbors.

Classifier

and

neighbors

to three. Okay. Then fitting the model.

I will write here classifier

dot fit

x train,

y train.

Okay. Then uh now we will predicting the

test uh test set results. Okay. So y

prediction equals to

classifier dot predict

x test. Cool. Okay. From Okay. Spelling

mistake.

Okay.

Okay.

Yeah. I guess it's fine now. So now uh

let's evaluate the prediction. So here I

will uh build the confusion matrix.

Okay. So here I will write cm confusion

matrix equals to confusion

matrix

uh y test

y prediction. Okay. Then here I will

write cm. Okay. So what is confusion

matrix basically? So I confusion matrix

uh you know is a table that shows how

well a model performs by comparing its

prediction to the actual values. Okay.

So what it show is a confusion uh matric

display the number of correct and

incorrect prediction of each class in

models whose either it can give you true

positive true negative false positive

false negative either it can give you

zero or one. Okay. So yeah moving

forward let's calculate the model

accuracy. This is the main part. If the

model accuracy is low it means uh your

analysis of you know classification is

not good. Okay. So accuracy it should be

more than 80 at least. accuracy

equals to

accuracy

score

y test

comma y prediction into 100 otherwise it

will give me in points so I will write

print

model

canon model accuracy

is

oh I will write plus STR I will round

it. Okay. Accuracy

I will add the person. So why I wrote

this accuracy to round? So I don't want

after points I need only two numbers.

Okay. I don't want like 6 7 8 9 10 11 12

like this. Okay. I need only like 80.20

like this. Cool. Let's run this. see

KN&N model accuracy is 96.67 67 that's

why I wrote two here and the percent

should be there so 96 which is very very

very good okay so this is how you can

find uh the accuracy so now let's find

the optimal number of neighbors in K

okay basically finding the best K so we

will use uh using cross validation

parameter okay tuning so first I will

create the list of K for KN okay so here

I will write K

list

equals to list

range

1 comma 50A 2.

Here I'm I will create the list of CV

score. Okay. So here I will write CV

scores.

Okay. Okay. I have to give brackets.

Yeah.

So we'll here perform the 10fold cross

validation. Okay, I will explain you

what is cross validation. Don't worry.

So first let me write for K N K listN

equals to

K neighbor classifier

and here I will add N neighbor

neighbors equals to K. Then scores

equals to cross

value score

KN&N then X train

Y train

okay

I will write here cross validation

equals to 10

comma scoring equals to I will write

here accuracy

Okay. And then here cv

course dotappend

to course dome.

Okay. So now yeah let me run it. Okay.

Comma some error came. Okay. The error

is the scoring parameter.

Yeah. Why? Because here accuracy you can

see and I did the spelling mistake.

Okay. So now what is cross validation?

Of course cross validation uh you know

determine the accuracy of your machine

learning model by partitioning the data

into two different groups. Okay called

training set and testing set. You can

see train test and the testing set. Okay

X and Y. So the data is randomly

separated into a certain number of

groups or subsets called folds. Okay you

can see the 10 folds we have wrote. Each

fold contains about the same amount of

okay and there is one more thing

validation. So validation is a technique

for assessing the accuracy of the model

on data set. Okay. And this cross

validation we did on new data set. Cool.

So now let's find the best K. Okay. So

here I will write best K equals to K

list. Okay. Before that I will write one

thing. I will write here MSE was

changing to mclassification error. Okay.

So equals to one

that's X for X in for X in

CV scores.

Okay. Okay, list I will write MSE dot

index

minimum

MSE.

I will use this bracket square bracket.

Okay. So then I will write print

the best

optimal

number of

neighbors

is person

best K. Okay, let's run it. See the best

optimal number of neighbors. Okay. N

neighbors is 9. Okay. So in the K

nearest neighbor KN algorithms. Okay. So

K represent the number of neighbors that

are considered when classifying a query

point. Okay. See the if we'll classify

this particular point you will get the

six. Okay. 1 2 3 4 5 6. Okay. And uh let

me show you. If you will classify this

portion only, so you will get 1 2 3 4 5

6 points like this. Okay. So the best

optimum number of neighbor is nine.

>> What is Python? Python is a high-level

object-oriented programming language

developed by Guido Van Roum in 1989 and

was first released in 1991.

Python is often called a batteries

included language due to its

comprehensive standard library. A fun

fact about Python is that the name

Python was actually taken from the

popular BBC comedy show of that time

Montipython's Flying Circus. Now let's

look at the top features of Python

first. So Python has a simple structure

and a clearly defined syntax. This

allows the learners to pick up the

language quickly. So it is easy to learn

and use.

Python can run on different operating

systems such as Windows, Linux, and Mac,

making it a portable language. It

enables programmers to develop the

software for several competing platforms

by writing a program only once.

Third, Python is freely available at the

official website since it is open

source. This means that source code is

also available to the public.

Now, Python uses an object-oriented

approach that encapsulates code within

objects.

Python provides a collection of

libraries for various tasks such as

machine learning, web development, and

data analysis. And finally, in Python,

you don't need to assign the data type

of the variable. When you assign some

value to the variable, it automatically

allocates the memory to the variable at

runtime.

Now, with that, let's move on to the

uses of Python programming.

So, Python programming language is used

to develop desktop applications and

build web applications too. It is

popularly used in the field of data

science, machine learning and artificial

intelligence to analyze data, build

predictive models and make business

decisions. Python is also widely used in

game development. Now, let's see some of

the popular Python frameworks and

libraries.

Python can be used for web development

using frameworks like Zango, Flask,

Pyramid and Churi.

Now you can build graphical user

interfaces using libraries and

frameworks such as Tkinter or just KER.

You can also use PI GTK, PIQT or PYJS or

Python JavaScript.

Now, Python is also used to perform

machine learning tasks using libraries

such as TensorFlow, PyTorch,

Scikitlearn, Mattplot Lib, and Scypi.

You can also perform mathematical

computations using numpy and pandas.

Now, let's look at the best ids that you

can use to write programs in Python and

perform specific tasks. So, we have

Jupyter notebook, which is part of the

Anaconda distribution that is widely

used these days. Even for our demo in

this video, we'll be using Jupyter

Notebook. I'll show you in a while. Then

we have the visual code editor from

Microsoft. This is also one of the

preferred IDEs by learners and

companies. Then we also have the popular

text editor called Sublime Text editor.

Then we also have PyCharm followed by

Python and Spider as our top idees. Now

let's look at the top companies that are

using Python in our day-to-day work.

So we have Google, Kora, Facebook, even

Netflix, Spotify, and Instagram. Now

there are other top product- based,

service- based and startups that also

use Python programming. So what really

is Python programming language?

Python is an object-oriented highle

programming language that supports

built-in data structures and dynamic

semantics.

It supports multiple programming

paradigms such as structured,

object-oriented and functional

programming.

Python is often described as batteries

included language because it has a

comprehensive collection of standard

libraries. Python supports different

modules and packages which allows

program modularity and code reuse.

Python was developed by Guido Van Rosum

and its implementation started in

December 1989.

Python 1.0 version was released in the

year 1994. Python 2.0 came out in

October 2000 while Python 3.0 was

released in December 2008.

Now that you have got an understanding

of the Python programming language,

let's now look at the top 10 reasons why

you should learn Python.

So at number 10, we have ease of use.

One of the most common reasons to like

Python is that it is quite easy to learn

and code. It provides a simple syntax

that improves readability and makes it

easier to understand. So developers can

create any desktop or machine based

application using this language. Python

is very versatile and is instrumental in

artificial intelligence and machine

learning. We will talk about this later

in the session. Compared to Java or C++,

it has fewer lines of codes.

In the example here, we are printing a

hello world program in Java. As you can

see,

if you have to write a program in Java,

you first have to declare the class name

along with its scope.

Next, using curly braces, you need to

pass the main method along with its

arguments. And then using

system.out.print print len method you

can print hello world that's quite a

tedious task isn't it

the same task of printing hello world

can be done using just one line of code

in python as shown here you can write

the print function and pass whatever you

want to display inside the brackets and

that will print the output it is so

simple

that is why Python is considered as a

highle language and it's open source

You can just download it from the

website and start using it.

At nine, we have active community.

You need a community to learn new

technology and friends are your best

asset when it comes to learning a

programming language. Python has large

community support.

It has an extensive and active community

to assist engineers, developers,

analysts, and data scientists with

expert support in case of programming

errors or issues with the software. You

can just go ahead and put your queries

in the community forum. The community

members will address your queries in

real quick time. Communities like Stack

Overflow also brings many Python experts

together to help learners.

Python enhancement proposals or PEP is

where the proposals and the improvements

are announced. Also, there are a set of

recommendations or core values called

the Zen of Python written by Tim Peters

that represents the guiding principles

for Python development.

Up next at 8, we have portable and

extensible.

Multiple cross- language operations can

be performed effectively because Python

is portable and extensible in nature.

For example, if the users have a Python

code written on Windows and they want to

execute on a Mac operating system or

Linux operating system or Solaris, they

can easily do it without any amendment.

They can also run this code on any

platform flawlessly and without any

interrupt.

Due to its extensibility feature, you

can integrate other programming

languages such as Java,Net, C and C++

codes with Python. The components of

other programming languages can be used

with Python and thus it can be used to

make a crossplatform suitable

application too. So it is a really good

feature that Python provides.

The next reason to learn Python is

testing frameworks.

Python supports several built-in

flawless testing tools and frameworks

that help in debugging and speeding of

workflows.

Some of the tools and frameworks

supported by Python are Piest, Selenium

and Splinter. This is the reason for

which every tester tries to use Python

based tools and frameworks to test any

application or code or to validate it in

an easier manner. Piest is the most

recommended testing framework for

functional, integrational, and unit

testing. You can run Selenium test

scripts using Python programming

language to automate various tasks. And

Splinter is an open-source tool for

testing web applications using Python.

It lets you automate browser actions

such as visiting URLs and interacting

with their items.

At number six, we have libraries and

packages.

Another reason why Python has become so

popular in the industry these days is

that it has a massive collection of

libraries and packages that make your

task simple and easy. It has a range of

libraries, packages, frameworks, and

modules for data manipulation,

statistical calculation, web

development, machine learning, and data

science.

Python programmers have developed tons

of free and open-source libraries that

you can use. You can find many of them

via Python package index, the repository

of Python software. Python provides the

default package called pip. Anaconda is

a third party Python ecosystem. Other

examples include numpy, sci and zango.

Then we have scripting and automation.

Python is not just a programming

language. It can also be used for

writing scripts for automating tasks and

workflows without human intervention.

The code can be written in the form of

scripts and executed later. Further, it

is interpreted by the machine and

checked for errors at runtime. The

machine is used to read and interpret

the code. Once the developer checks the

code, it can further run or be used

several times without any interruption.

This allows you to automate a set of

certain tasks within a program or the

same code can be used with other

applications as well.

At number four, we have web development.

Another reason to learn Python is that

it makes the web development process so

much easier.

It provides a wide collection of

frameworks that make it easier for

developers to develop web applications.

Some of the examples are Zango, Flask,

Pyramid, Turbo Gears, CherryPie, etc.

These frameworks are written in Python

which makes the code a lot faster and

stable.

The task which used to take hours in PHP

can be finished in minutes using Python.

Python is also used for web scraping.

Django offers many elements of intricate

programs such as template design,

management panel, signing in, signing

up, signing out, URL routing, etc.

Once the user establishes the framework,

all these features become ready to use.

Flask is a microwave framework written

in Python.

of all the components that are part of

this module, they are all ready to

execute in the server context.

Pinterest and LinkedIn use Flask.

Pyramid offers more attributes than

Flask. It will assist users with URL

routing and authentication support.

Turbo Gears is a highly recommended and

scalable framework that supports

features such as authentication,

caching, identification, management of

sessions, and pluggable applications.

Up next at number three, we have machine

learning.

The growth of machine learning has been

phenomenal in the last 5 years and it's

rapidly changing the world around us.

Python is one of the most preferred

programming languages for machine

learning because of its simple syntax

and support for several machine learning

libraries.

Using different libraries and functions

in Python, the system can learn and

train itself from past data.

Once the system is trained, it can then

learn to adjust itself to new inputs.

Finally, it can make predictions and

perform humanlike tasks automatically.

At number two, we have data science.

Machine learning and data science go

hand in hand. Python is robust, scalable

and provides extensible visualization

and graphics options. Hence, it is

widely used in data science.

Python has libraries such as numpy for

numerical computation of data, pandas

for operations to manipulate data on

numerical tables and time series. It

also provides simply for symbolic

computation and sci for technical and

scientific computations.

It has another library called pyrain

which is sought for python based

reinforcement learning, artificial

intelligence and neural network library.

Scikitlearn is the machine learning

library for creating classification,

regression and clustering algorithms.

And finally, it provides PyTorch and

TensorFlow for deep learning.

Finally coming to the most important and

the top reason to learn Python which is

career opportunities and salary.

Python language provides a variety of

job opportunities and promises a high

growth graph with huge salary prospects.

It is been used by most of the tech

giants.

Industry leaders using Python are

Amazon, Google, Facebook, IBM, NASA,

Netflix and YouTube.

Next, you can see the Google trends

but I have considered three programming

languages Python, Java and C++. I have

compared them for the past 12 months.

You can see it clearly on your screens

that Python has become a frontr runner

in terms of popularity and web search

volume. It means people are interested

in Python. They want to learn it and use

it in their work. You can also check for

the YouTube search.

There also you will find that Python

programming language is the most

searched language on YouTube.

Now on your screens you can see the

report of PPL which is popularity of

programming language index. It is

created by analyzing how often language

tutorials are searched on Google. It is

a leading indicator. The raw data comes

from Google trends. The bar graph

depicts that Python is the most popular

and widely used programming language

across the globe followed by Java then

JavaScript and C.

The popularity of programming language

index can help you decide which language

to study or which one to use in a new

software project.

The next graph shows the popularity of

Python and Java over the years starting

from 2004 till the current period which

is 2020. Worldwide, Python is the most

popular language. Python grew the most

in the last 5 years by 19.4%. 4% and

Java lost the most by minus 7.2%.

Now let's talk about the different

career opportunities and the job roles

that you can get into if you learn

Python language.

First, you can become a Python developer

where you will be asked to write and

test codes, debug programs, and

integrate applications with third party

web services.

Second, you can become a web developer.

Here you will be responsible for writing

serverside web application logic. Python

web developers usually develop back-end

components, connect the application with

third party services and support the

front-end developers by integrating

their work with the Python application.

You can also become a data analyst if

you know Python. As a data analyst, you

have to gather data from multiple

sources using scripts. analyze that

data, develop and implement databases

and data collection systems.

You can become a data scientist. As a

data scientist, you need to understand

the challenges in business and come up

with the best solutions using modern

tools and techniques to analyze,

visualize, and build prediction models

to make business decisions.

Lastly, you can be a machine learning

engineer where you can develop

intelligent machines that can learn from

vast volumes of data and apply knowledge

without human intervention.

So there's a lot of scopes if you learn

Python. But before we move on, let's

understand first what is Jupyter

Notebook. So guys, as you can see all

over here that Jupyter Notebook is a

popular open-source tool that basically

allows you to create and share documents

which contains codes, equations, you can

have visualizations also. Basically, it

is used for data analysis, machine

learning and scientific research which

makes it a very essential tools for

developers like data scientists and

researchers alike. Now before installing

Jupyter notebook I request you that you

have Python installed in your system. So

the requirement should be Python 3.6 or

greater. So now let us officially

navigate to the Python's website. So

guys as you can see all over here. So on

python.org if I click on download

Python. So we're going to see that all

over here download Python 3.125. So as I

already told you that the requirement of

Python should be greater than 3.6. So

just you can click all over here and you

can see the download has started.

So guys as you can see all over here

that we have installed the Python. Now

let us open the file. So you can see the

given software is going to installed on

this directory. Okay. So just click all

over here. So guys as you can see all

over here the Python installation of

3.125 is in progress. Let's wait for

some time till it gets installed.

So as you can see guys all over here

that we have successfully installed our

Python. Now let us open our terminal and

let us check whether Python is correctly

installed. So we are going to type

python

/ version.

So as you can see all over here we have

successfully installed our Python. So

guys that was our prerequisite. Now

there are two ways to install Jupyter

notebook. The first one can be pip.

Okay, pip is a package manager or using

Anocanda distribution. So let us see

with pip first. So guys, pip is a

package manager which is used to install

and manage software packages libraries

written in Python. So you can see all

over here that the Python with version

greater than 3.6 have default pip

installed in them. Okay. So we can use

pip command to install our Jupyter

notebook. So guys as you can see all

over here we have come to the official

documentation of jupitter.org and it is

saying that installing Jupyter lab with

pip command. So what you can do guys you

can just copy all over here. You can go

right all over here and click on this.

Now as you can see all over here it has

started downloading the Jupyter lab.

So guys, we are going to install our

Jupyter lab with the pip command. So

this is the official documentation of

Jupyter notebook. Okay? And just all you

have to do is copy this and type on your

terminal. So as you can see all over

here it has started downloading the

packages which is required to download

the Jupyter notebook. Let us wait for

some time.

Okay guys, so we have successfully

completed this step. Now let us move on

to our next step. So as you can see all

over here. So we have installed. Okay.

Then what we have to do then you can

type this. We can launch the Jupyter lab

with this command on the terminal. Now

let us wait. So as you can see all over

here guys, we have successfully

installed our Jupyter notebook. So you

can go all over here and just create a

new notebook and you can also choose

your kernel and you can start working on

your Jupyter notebook. Suppose I'll show

you one snippet. So 3 + 5. Let us try to

run this notebook. So as you can see it

is giving us the eight as answer. So it

is following the Python syntax and in

this way we have successfully installed

our Jupyter notebook using the pip

command. So now as you can also see all

over here you can also install Jupyter

notebook with this command pip install

notebook and then you can just open it.

This is also an another alternative.

Similarly, you can install with VA also

same command and just open the VA. Now,

if you are using any other operating

system like Mac OS or Linux, then you

can install by brew install Jupyter Lab.

So, home will be the package manager for

Mac OS and Linux. So, I hope so you are

pretty clear with how to install Jupyter

notebook with the pep command. Now, I

have downloaded Anacondas from this

official website. So as you can see all

over here this is the official website

of Anaconda. Okay. Now just type your

email and you can just download it. So

similarly as you can see after

installing I'm going to launch my

installer and let us click next. Okay.

Let us click agree. Okay. And let us

install this on the given directory.

Let us wait for some time till the

installation gets complete.

So guys as you can see all over here we

have completed our installation of

Anoconda. So just click on finish and

you can say we have successfully

installed our Anocanda. Now let us open

our Anaconda navigator. So just click

on.

So as you can see all over here just

right click on this and our Anocanda

navigator will be opened. So as you can

see all over here this is our Anocanda

navigator and it is loading the packages

and for us to install the Jupyter

notebook. So as you can see all over

here just click on launch. So guys if

you click on launch it is going to open

our Jupyter notebook. So as you can see

all over here it is saying launching the

Jupyter notebook and it is hosted on

localhost 8889. So this is our hosted

Jupyter notebook and in similarly you

can create a new notebook all over here

and in this way you can start working

>> LLMs. If you ever wondered how machine

learning can now understand and generate

humanlike text, you are in the right

place. From chatboards like Chat GPT to

AI assistant that powers search engines,

LLMs are transforming how we interact

with technology. One of the most

exciting advancement in this space is

Google's Gemini or OpenAI Charging large

language model designed to push the

boundaries of what AI can achieve. In

this video, we will explore what LLMs

are, how they work, and why models like

Geminy are critical for the future of

AI. Google Gemini is part of a new wave

of AI models that are smarter, faster,

and more efficient. It is designed to

understand context better, offer more

accurate responses and integrate deeply

into service like Google search and

Google Assistant, providing more

humanlike interactions. So we will break

down the science behind LLMs including

their massive training data set,

transformer architecture and how models

like Gemini use deep learning innovation

to change industries. Plus we will

compare Google Gemini to other popular

LMS such as OpenAI Chity models showing

how each of these technologies is used

to power chat bots, virtual assistants

and other AIdriven application. By end

of this video, you will have a clear

understanding of how large language

models like Gemini work, their key

features, and what they mean for their

future AI. Don't forget to like,

subscribe, and hit the bell icon to

never miss any update from Simply Learn.

So, what are the large language models?

Large language models like Chargen

pre-trained transformer 4 o and Google

Gemini are sophisticated AI system

designed to comprehend and generate

humanlike text. These models are built

using deep learning techniques and are

trained on vast data set collected from

the internet. They leverage self

attention mechanism to analyze

relationship between words or tokens

allowing them to capture context and

produce coherent relevant responses.

LLMs have significant application

including powering virtual assistant

chatboards, content creation, language

translation and supporting research and

decision making. Their ability to

generate fluent and contextually

appropriate text has advanced natural

language processing and improved human

computer interaction. So now let's see

what are large language model used for.

Large language models are utilized in

scenarios with limited or no domain

specific data available for training.

These scenarios include both few short

and zero short training approaches which

rely on the model's strong inductive

bias and its capability to derive

meaningful representation from a small

amount of data or even no data at all.

So now let's see how are large language

models trained. Large language models

typically undergo pre-training on a

board. All encompassing data set that

shares statical similarities with the

data set specific to the target task.

The objective of pre-training is to

enable the model to require highlevel

feature that can later be applied during

the finetuning phase for specific task.

So there are some training processes of

LLM which involves several steps. The

first one is text prep-processing. The

textual data is transformed into a

numerical representation that the LLM

model can effectively process. This

conversion may be involve techniques

like tokenization encoding and creating

input sequences. The second one is

random parameter initialization. The

model's parameter are initialized

randomly before the training process

begins. The third one is input numerical

data. The numerical representation of

the text data is fed into the model of

processing. The model's architecture

typically based on transformers allows

it to capture the conceptual

relationship between the words or tokens

in the next. The fourth one is loss

function calculation. A loss function

calculation measures the discrepancy

between the model's prediction and the

actual next word or token in a syntax.

The LLM model aims to minimize this loss

during training. The fifth one is

parameter optimization. The model's

parameter are registered through

optimization technique. This involves

calculating gradient and updating the

parameters accordingly gradually

improving the model's performance. The

last one is iterative training. The

training process is repeated over

multiple iteration or epox until the

model's output achieve a satisfactory

level of accuracy on that given task or

data set. By following this training

process, large language model learn to

capture linguistic patterns, understand

context and generate coherent responses

enabling them to excel at various

language related tasks. The next topic

is how do large language models work. So

large language models leverage deep

neural network to generate output based

on patterns learned from the training

data. Typically a large language model

adopts a transformer architecture which

enables the model to identify

relationship between words in a sentence

irrespective of their position in the

sequence. In contrast to RNAs that rely

on recurrence to capture token

relationship transformer neural network

employ self attention as their primary

mechanism. Self attention calculates

attention scores that determine the

importance of each token with respect to

the other token in the text sequence

facilitating the modeling of intricate

relationship within the data. Next,

let's see application of large language

models. Large language models have a

wide range of application across various

domains. So here are some notable

application. The first one is natural

language processing NLP. Large language

models are used to improve natural

language understanding tasks such as

sentiment analysis, named entity

recognition, text classification, and

language modeling. The second one is

chatbot and virtual assistant. Large

language models power conversational

agents, chatbots, and virtual assistant

providing more interactive and humanlike

user interaction. The third one is

machine translation. Large language

models have been used for automatic

language translation enabling text

translation between different languages

with improved accuracy. The fourth one

is sentiment analysis. LLMs can analyze

and classify the sentiment or emotion

expressed in a piece of text which is

valuable for market research, brand

monitoring and social media analysis.

The fifth one is content recommendation.

These models can be employed to provide

personalized content recommendations

enhancing user experience and engagement

on platforms such as news website or the

streaming services. So these application

highlight the potential impact of large

language models in various domains for

improving language understanding

automation. So hello guys welcome to

this demo part of this video. So here

what I will do I will go to new then

Python 3 file

then here

I will give it the name called

exploratory

data

analysis. Basically we will so we have

one data set file of

roller coaster basically. So we will be

using that and you can download that

file from the description box below from

the below link driving link. Okay. So we

will be doing some small basic functions

using Python and later on we will uh

make some good charts. Okay. We will

remove duplicates and all we will do all

that

thing. We'll do data preparation. We'll

do feature engineering. Okay. And uh

we'll remove the duplicates. We'll check

for the duplicates. We'll make charts

like histogram, KD blocks, box plot and

like many more similar to that like heat

map. We will make scatter plot group by

comparison. Okay, we all do that, right?

So just stick with me and you can write

side along with me here this code. Okay.

Okay. So let's start with importing

pandas first.

I guess everyone know what is pandas and

numpy

fine

as np then I will import

numpy

as np why this is np and pd sorry my bad

pd is here because I don't want to write

again and again this pandas this numpy

so basically I can write this small

version okay Yeah. So if you guys don't

know what is panda. So panda is very

popular library for working with data.

Okay. It's goal is to be the most

powerful and flexible opensource tool.

Okay. So it has reached that goal. So

data frames are the center of pandas.

And what is data frame? A data frame is

structured like a table or a

spreadsheet. Okay. The rows and the

columns. Okay. Whereas numpy numpy is an

open-source Python library again that

facilitates uh you know efficient

numerical operation on large quantities

of data. Okay. So there are many

functions in numpy as well those we can

use in the pandas data frame. Got it. So

we'll import one more

dot piplot

dotp. Okay. as

plt. Okay, there is one more library

mattplot lip for plotting the graph and

and there is one more import

seabbon

as SNS. So what is se? Sebon is again

the Python data visualization library

based on Matt plot lip. Okay, it

provides a highle interface for drawing

attractive and informative statical you

know graphs. Okay. Then I will write

here plt dot style

dot use.

We'll write ggplot.

Got it. ggplot.

Fine.

So yeah,

let me run it. So what is ggplot? So

ggplot is an again this is an opensource

data visualization package for

historical programming. Okay. So yeah uh

you can say or a general scheme for data

visualization which breaks up graph into

you know semantic components such as

scales and layers. Okay. ML lip py

pyab

p. Okay. Yeah. So now

we will import our data set df. DF means

data frame. You can write any word of

your choice. Then pd again pandas dot

read

csv used for readings CSV file. Okay.

Excel files. Got it? Then here I will

write my

this coaster

dot CSV. Okay. I'm not writing any path

because my this uh data set is here

itself. Coaster Coaster CO this is

coaster.csv CSV. Okay. If you have your

data set in another location as in like

in C drive, D drive or whatever. Okay.

You can give that path.

Okay. Let me run it. Yeah. So now we

will do some data understanding. Okay.

We'll see data frame shape head and tail

data types and describe like small small

function. We will use DF dot shape.

Right? So we have 1087 rows and 56

columns in our data set. Okay. Then we

will see df.hat

five.

So df do.head means it will give me top

five rows of my data set. Okay. You can

see 1 2 3 4 5. Okay. Five rows and 56

columns. Here 56 column but we have 1087

rows. Okay. So we have coaster name,

length, speed, location, status, opening

date, type, this, this, this, this.

Okay, we will do

uh you know we'll make some graphs using

this these columns. Okay, and there is

one more df.tail

again last five rows. So df.tail gives

you the last last five rows. See 1086

1085. Okay. And if you want to see

full data, it is here. Okay. 0 to 108.

Fine. Yeah. So we have one more df doc

columns to check

all the columns.

See coaster name, land, the speed,

location, status, opening date, type and

these all are my columns names. Fine. So

we have one more data types. Actually we

have like many data types but let me

show you some important ones or you can

say some the basics on one. Okay. So DF

dot D types. Dypes means data types.

Coaster name is object type length

object and inversion is float. Okay.

We'll check inversion.

Where is inversion? Yeah float type.

Okay,

then everything is object and eer

introduces int numeric one and latitude

float. Okay, so basically we have three

types float, int and object.

Okay, then what we will do? Let's just

quickly check this describe.

Okay. So what is the count of this

inversion 932?

Okay. It won't include any you know what

is it empty cells. Okay. It won't count

empty cells. Right. Mean of this

particular table then standard deviation

then mean minimum value 25% 70% 75% and

max. It will describe you all this.

Okay. And there is one more df.info info

to get the info. See coaster name 1087

value non null object then length 953

null. Okay. Like this.

Yeah. So now moving forward what we will

do? We will do data preparation. Okay.

What comes in data preparation like

dropping irrelevant columns and rows

which we don't want. Okay. Then second

thing is identifying duplicates columns.

Then third is renaming columns.

Then we'll do some feature creations,

right?

So if you want to drop a column, okay,

how you can drop?

So you have to write just df dot drop.

First I will give here stag data

repage.

Okay. Yeah. So how you can drop a

column? Okay. DF.

Then here I will write

opening

date=

to 1.

Okay.

Opening. Okay. X is wrong.

Instead of this

maybe what is the opening date? Okay. O

is capital here.

Instead of this you can use double

equals to

okay axis is not defined.

My bad. So what you have to do? You have

to give this here and yeah.

Okay. Next is not defined.

Yeah. Fine. Okay. So as you can see here

first I will show you this question name

length speed location status and opening

date is there fine.

So if you will go here question name

length speed location status nothing

opening date is there right so this is

how you can drop a table fine for a

while I'm making this as a comment maybe

in future

upcoming you know making graph I'll

leave this okay so yeah so I will write

here df equals to

df

poster name,

comma, location

then comma

status.

Then here I will write manufacturer.

Fine. Then again comma.

Then I will give here.

Okay. Year

introduced

then

comma

latitude.

Wait I will tell you why I'm doing this.

Then longitude

latitude longitude

then

type main.

Okay. Then

opening

date

clean.

Fine.

Then speed into

m/ hour r

speed

and comma what else then

height

foot

inversion

clean

geforce

clean dot copy. Yeah.

So here I will write okay

type in

okay one more mistake is here

fine.

So now what I will do

I will write here opening

date

clean equals to PD2

dot date time

DF

opening

date

Clean fine.

What happened?

So what does this pd do to date time do?

Okay. So it converts argument to data

time date time. Okay. So this function

you know you can say converts a scalar

array like or series or data frame

dictionary like to a pandas datetime

object. Okay.

So now let's rename the columns.

Okay.

Then to you know for better things DF

equals to DF dot rename

columns equals to

poster

name. Then

poster

name,

year

introduced then

here

introduced.

Okay.

Opening date clean to

open

date.

Okay. Then I will write speed

I then write caps speed.

Got it? Then

height

into foot

I will write it as

catch 50.

So why I'm writing this because this is

a very good practice as a professional

way or as a data analyst or as a

business analyst whatever you are

working on machine learning projects or

whatever this is a good practice.

Okay. So inverions

in versions

clean

then should be like

invers

fine.

Let me run it.

Okay. Now let's check the column.

Yeah. So now you can see our column name

is changed.

Fine.

So I will write this na

dot sum.

So what does this is na dot sum do? So

is na function returns a boolean value

of you know true if the value is n and

the false otherwise and the sum function

returns the sum of the true values which

equals to the number of n values in the

column. So here we have zero and n here

same the status 213

here 217 5

height foot is 916 okay so now what we

will do we will write here df dot

location

df dot duplicate

I told you in starting

we'll

duplicate

Okay.

So, question name, location, status,

man. Okay. It's duplicated now. Okay.

So, now what we will do? I will check

duplicates for the coaster name. So, you

can write df. Loc

then df

df dot

duplicated and subset equals to

poster

name then dot head. Okay. Head of five.

Now everyone know right what do

what this head does.

Okay. Yeah. So here you can see so these

are duplicate okay of the question name

right.

Why? because you can just check the

typing

and all. Fine. So now checking with an

example duplicate. Let's check

dot query

question name

crystal.

Now you will get each equals cyclone.

Okay. I took this crystal B cipher.

Fine.

Then run it. So now you can see 39 and

43 are the same.

Okay. Everything is same. So this one is

duplicate. So now what I will do? I will

write here. Just let me give some space.

Yeah. DF dot columns.

Then I will write here DF

dot location.

Then TF dot

duplicated

subset

question

name

then

location

then

opening

date.

Okay.

dot

d set

index

then drop that.

Okay.

So now what we will do we will do some

feature understanding.

Okay.

So now we will do some feature

understanding.

Okay.

So in this we will plot some feature

distribution like histogram KD box plot.

Okay. For that I will write here DF

year introduced.

Okay. Then I will write value

count.

So now what I will do? Let's create bar

chart. Okay.

Ax goals to der

introduced dot

value

counts. Okay. Then dot

add 10 max. I need

then dot plot kind equals to I need bar

then title is top

I write what what what should we give

top 10

years

coasters introduced.

Okay, then

I will write here ex dot set

X label. Then I will write here

introduced.

Okay. Then I will write here ex dot

set Y label.

than

count. Okay, let me run it. Okay,

unexpected character after line

continuation.

Okay,

we can remove this

unexpected intended.

Now let me run it. Yeah, so this is our

bar plot. Okay. So this is how you can

create bar plot using your data. Okay.

So a bar plot is you know one of the

most common types of graphics in this

data visualization or EDA. It shows the

relationship between a numerical and the

categorical variable. Okay. So each

entity of this categoric variable is

represented as a bar.

Got it? So now let's do some more. So

what uh let's make a stogram class.

Okay,

histogram.

So for that I will write

a ex = to df

speed

r then plot

kind equals to

Then comma I will write bins here. Okay.

Bins equals to 20

then

title

coaster

speed.

Okay. meter per hour then AX

dot set

X level

speed

okay forgot to give this

okay yeah so histogram displays

numerical data by grouping data into

bins of equal width so here each bin is

plotted as a bar whose height correspond

to how How many data points are in that

bin? So bins you can say are also

sometimes called the intervals or

classes or buckets. Basically

it's the same.

So here for this this much is the bin.

For this this much is a bin. This is

like that. Okay.

So now let's create KD plot. So for that

it's very simple. Ax = to DF.

Then I can write as speed in me of R

then plot

kinda

then title

I can write the coaster

speed

then ex dot set

X

speed

can.

Yeah. So, KD, what does KD means? A

kernel density estimate. This plot is a

method of, you know, visualization the

distribution of observation in a data

set analog to a stola. So, KD represent

the data using a continuous probability

density curve in one or more dimension.

Okay. In one or more dimension that

curve fine. So now let's do some feature

relationship. So in this we will make

scatter plot, heat map correlation and

pair plot. Okay. Or we can also do some

group by comparison. Fine. So now let's

first make

scatter plot. Okay. So, df dot plot

kind equals to

scatter

comma

x = to

speed

m/ hour

comma y = to

height in foot.

Fine.

Then comma let's give the title

equals to coaster speed

versus

height.

Fine then pl do

okay some error is there

speed meter per hour. Okay. Okay. S

should be capital and H should be

capital.

Yeah.

So this is scatter plot speed versus

height. This is a speed versus height.

Okay. If the I can see here height is

directly proportional to speed somewhat

because if you can see the 350 is the

speed less than

120 km/h. No it's not like that. Okay

fine.

So a scatter plot identifies the

possible relationship between change

observed in two different sets of

variable. Here the variables are height

and the speed. Okay. It provides a

visual and aesthetical uh you know means

to test the strength of relationship

between two variable. Fine. Now let's

okay now let's make one more to give you

better idea. Okay with legend.

Now I will wait SNS dot scatter

plot then x = to

speed

r

y = to

height into foot

whatever you can say then hue will be

there the year

introduced.

Okay. Then data equals to DF.

Then ex= to set title.

Set title. Then

coaster

speed versus

height.

then plt dot show.

Yeah.

So now you can see

so this color you know the light color I

will let me zoom it. Okay. So this color

we have some here which are introduced

in '90s. Okay. And this color which are

introduced in 1925 and these color which

are introduced in 2000. Okay. So you can

see like this as well.

So now let's make pair plot. Okay, it's

look amazing.

So let me write SNS dot

dot pair plot then TF comma

variable is equals to I need

introduce

then speed

meter per R

comma S should be capital height

foot. Just remember the spelling, okay?

It's case sensitive.

Inversion

inversions, comma,

inversions, comma, geforce.

Okay. Comma. U I will put what? What?

What? What? Okay. I will put type

main.

Fine then pl do

okay let me run it okay some error is

there

year introduced

y capital

let's say here so that's why I'm saying

just remember the proper spelling again

some

error in versions is version.

Now let me run it.

Okay. Again some

what inversions?

Okay. Let me check here while renaming

inversions only.

Okay. Fine.

in versions.

Let me paste in but nothing changes. Let

me check again.

Okay, some error is there.

Wait. So yeah, you can see it's running

fine.

So after this what I will do?

So what is the pair plot? So pair plot

function allows the user to get you know

then an axis grid via which each

numerical variable is stored in the data

is shared across the x and the y okay in

the structured column.

So this is how see you know type mean

wood other and steel this red one is

wood other or the blue one and the

purple are the steel one okay the

different different format okay now

let's create the last graph which is

heat map okay so let me write df

correlation equals to df

here introduced,

comma,

speed m/a

height

into foot

canions,

comma

G force

then

drop then coalition okay DFO

introduced I

capital yeah

so So this is correlation values of all

the things right. So now let's write SNS

dot

heat map

heat map TF C

not will be true.

Okay.

Now let me run it. Yeah. So to create

heat map in Python so you can use this.

Okay. C bond library for this heat map.

So this function takes a data frame you

know as a input and generates a heat map

type of things as the output. Okay. So

this is how you can perform EDA using

any data set or you can show your data

or insights with a beautiful

representation using graphs and all like

this. Fine. Web scraping is a powerful

technique that allows you to

automatically extract data from website.

Turning the vast amount of information

available online into something you can

easily analyze and use. Whether you are

gathering data for research, building a

data set for machine learning project,

or just curious about how websites work

behind the scenes. Web scraping is an

essential skills to have in your

toolkit. On the other hand, Python is

one of the most popular programming

languages for web scraping thanks to its

simplicity and the wealth of libraries

available. In this video, we will

explore how to use Python to scrape data

from website. And we will dive into

practical examples using Python

libraries like request and beautiful

soup to fetch and parse web content. But

it's not just about the code. Web

scripping comes with its own set of

challenges and ethical constitution. We

will talk about how to scrape

responsibly, respecting the rules set by

websites and ensuring that your scraping

activity don't negatively impact the

sites you are collecting data from. So

by the end of this video, you will have

a solid understanding of how to start

scraping data from the web using Python.

Whether you are new to programming or

looking to add web scraping to your

skill set, this video will give you the

knowledge and tools you need to get

started. So let's jump in and see how

Python can help you unlock the full

potential of the web. So without any

further ado, let's get started. So here

I am using this Google Collab for the

web scraping. Okay. You can use your own

like Jupyter notebook, Visual Code

Studio, any thing. Okay. So here I'll

write scraping

using Python.

Okay. Then here first you have to

install some libraries like

you know uh request

and you have to install beautiful. So

you have to install that you know pandas

because we will create one data frame

and we will save it then we will check

our data. Okay. And you can install some

basic Python library like numpy and all

that. Okay. So here first I will import

request.

Okay. Then I will write from PS4

import

beautiful

soap.

Okay.

So

then I will write import

pandas

as pay.

Fine. Then now what I will do? Okay

let's see what is this request and all

the so request is an HTTP client library

for the Python programming language. So

request is one of the most you know

downloaded Python libraries. Okay. It's

like more like over 2,000 not exactly

2,000 sorry 200 or 300 million monthly

download. Okay. So what it does it maps

the HTTP protocol onto Python subject

oriented semantics. And here beautiful

soap. So

SOAP is a Python package for you know

parsing the HTML and XML documents

including those with you know malphone

markup. It creates a parse tree for

documents that can be you know used to

extract data HTML from HTML which is

useful for web scraping. Let me run it

here. Our second step will be

define the URL and the headers. Okay,

URL and the headers.

So URL okay from where you want to

extract your data. So I will write here

simply learn

then this okay let's open this PMP

certification

it yeah so here I will write this

and headers

equals to

so these are the headers okay so like in

which you know device or you're working

on which browser you are working on. So

these are for the headers. So now what

what I will do I will send a get request

to the simple page. Okay to this URL for

that you know right let me write this

sending

get request.

Okay, here write response

equals to

request

dot get

then URL

comma headers

equals to headers.

Okay, let me run it. Okay, working fine.

So now uh let's check if the request is

successful or not. For that if

response

dot status code plus equals to = 200

then

so equals to

false.

Okay. Then response

dot content

do

HTML

dot parser. Okay.

Here.

So here I will write course

titles. Why? because here I'm you know

initializing the list to store our

titles or whatever the things are okay

basically the data okay why I'm writing

course here because this is again course

page that's why nothing else okay here I

will give the empty list

so now we will find all course titles

based on the actual HTML structure what

is HTML structure

you have to go to the page right click

Click then inspect.

Okay. According to this HTML structure

means what is the class name? What is

the you know this PMP certification is

which heading? H1 heading, H2 heading,

which heading it is. Okay. So let's see

which heading it is. Okay. Okay. This

PMP certification training H1 heading.

Okay. Just remember H1 heading.

So here I will write for

quotes in soap

dot find

or

h1.

Okay.

Then title

custom course text

dot strip.

Okay.

Then here I'm write course

titles

dot append

title. Okay.

Now what I will do? I will check

if any data was extracted before it not.

Okay, I will write if course

titles.

Here I will create a data frame. DF

equals to PD dot data frame.

In this I will write course title. it

will be our you know uh that column

name. So for I will write it course

data. Okay.

Then

course

titles.

Okay.

Then here I will write df

dot to

csv.

Then uh just create simply learn dot

CSV. Okay. So our data will be saved in

this simply learn dot csv. Okay. CSV.

Fine. Here what I will do I will write

index equals to false.

Print

data.

Print data is saved in CSV.

Simply learn dot CSV.

Fine.

So here I will write

else

print

web page.

Okay.

write field

to

retrieve the web page. Okay, that's it.

Let me run it.

Okay, here if response storage to this

group

spawn object has no attribute

status code. Okay.

Okay. Title.

It is text.

Some minor spelling mistakes are there.

Okay. Dear.

Sorry. Sorry. My bad

again.

Okay. Sorry.

D will be capital, F will be capital.

Yeah. So now you can see data is saved

in simply dot CSV. So what I will do? I

will write data equals to everyone know

how to read CSV in Python. Read

CSV.

What was our file name? Simply learn

CSV.

Okay. Let me copy paste. Run it. Link

fine data.

Okay.

PMP certification training data is fed.

Okay. Here you can see H1 we gave and

that's why PMP certification training

came. Let's take what's in H2.

Okay. In H2 there is leading premier PMI

partner something is there. Okay. I will

change here

H1 to H2 then I will run it. Okay

fine.

See leading premier P P P P P P P P P P

P P P P P P P P P P PMI. Okay, let me

make it bigger. Yeah. So now you can see

course data we mentioned. So

if you will see this H1 didn't come why

because we have mentioned H2 that's why.

So these all are the H2.

Okay. So what is this? I don't know.

Okay. This is a time series graph. This

is Google collaping. Okay. So this is

how you can literally retrieve your data

from any website. Okay. Just remember

that some websites don't give you access

to their you know for their scrapping

just like Amazon don't give so you have

to use at that time API and this is not

ethical also to use someone's data okay

without asking or whatever you can

>> welcome to the neural network tutorial

my name is Richard Kersner I'm with the

simply learn team what's in it for you

well today we're going to cover what is

a neural network what can neural neural

networks do, how does a neural network

work, types of neural networks, and then

we're going to jump into a use case to

classify between the photos of dogs and

cats, and we'll do that on the KAS with

the TensorFlow in the back, but it's a

Python script. So, that's always my

favorite part is when we dive into the

actual script. So, what is a neural

network? So, hi guys. I heard you want

to know what a neural network is. Here

we have uh looks like he just went

shopping at a red tag sale. My robots's

back. So as a matter of fact, you have

been using neural network on a daily

basis. In today's world, it's just

amazing how much we use our new

technology. We're not even aware of it.

When you ask your mobile assistant to

perform a search for you, you know, like

saying you're Google or Siri or whoever

you use, Amazon Web, self-driving cars.

So that's the newest thing coming out.

They're just now trying to make those

legal in different states in the US and

around the world. Even in the UK, they

now have self-driving cars going up and

down the street. It's pretty amazing.

These are all neural network driven.

Computer games use it. A lot of computer

games are driven by neural networks in

the back end as part of the game system

and how it adjusts to the players. And

it's also used in processing the map

images on your phone. So every time you

do a navigation someplace and it opens

it up, they now use neural networks to

help you find the quickest way to get

there. Neural network. A neural network

is a system or hardware that is designed

to operate like a human brain. In

today's development, this is so

important to understand because we don't

have anything else to compare it to. I'm

sure someday in the future, the computer

will redefine or the neural network or

the AI artificial intelligence will

redefine what these mean. But as far as

we can today's world, in today's

commercial development, we have to

compare it to what humans do. So, it's

we want to compare and how it operates

to a human brain and how it solves

problems like a human does. What can a

neural network do? And really, we're

just going to dive in deeper to we just

covered and look at other examples. So,

what can a neural network do? Well,

let's list out the things neural

networks can do for you. Translate text.

Boy, we got Google Translate and

Microsoft has their own translate. They

have some really cool. They actually

have an earpiece. It's supposed to start

translating as you talk. What a cool

technology. What a cool time to live.

Identify faces. Can you imagine all the

uses for facial identification? In the

case of our uh sample or our code that

we're going to look at later, we'll be

identifying dogs and cats. So, not quite

as detailed as uh understanding whose

face belongs to who. I I'm waiting for

the Google glasses to come out so I can

see who's who and the identify faces as

I'm walking around. Have a little name

tag over them. Not out there yet, but

boy, we are close. We can identify the

faces and they have all kinds of

technologies to bring that information

back to us. Recognize speech goes along

with the translate text. So now as

you're talking into your assistant, it

can use that to do commands, turn lights

on, all kinds of things you can do with

recognizing speech. Read handwritten

text. They're starting to translate all

these old text documents that they've

had in storage instead of doing it

individually where somebody's going

through each text by themsel in a room.

Picture like an old Raiders of the Lost

Arc theme where he's in the back, you

know, archaeologist studying the text.

Now it's fed into a computer. They take

a picture. They even use neural networks

to take a scroll that is so messed up

that they can't undo the scroll and they

x-ray it and then they use that x-ray to

translate the text off of it without

ever opening the scroll. I mean just way

cool stuff they're starting to do with

all this. And of course control robots.

What would be a neural network without

bringing in the robots? And we have our

own favorite robot in the middle who

goes to our red tag cell and goes

shopping for us. So, you know, these are

just a few of the wonderful things that

neural networks are being applied to.

It's such an infant stage technology.

What a wonderful time to jump in. And

there are a lot of other things it goes

into. I mean, we could spend just

forever talking about all the different

applications from business to whatever

you can even imagine. They're now

applying neural networks to help us

understand. So, now we talked a little

bit about all the cool things you can do

with a neural network. Let's dive in and

say, how does a neural network work? So

now we've come far enough to understand

how neural network works. Let's go ahead

and walk through this in a nice

graphical representation. They usually

describe a neural network as having

different layers. And you'll see that

we've identified a green layer, an

orange layer, and a red layer. The green

layer is the input. So you have your

data coming in. It picks up the input

signals and passes them to the next

layer. The next layer does all kinds of

calculations and feature extraction.

It's called the hidden layer. A lot of

times there's more than one hidden

layer. We're only showing one in this uh

picture, but we'll show you how it looks

like in a more detail in a little bit.

And then finally, we have an output

layer. This layer delivers the final

result. So the only two things we see is

the input layer and the output layer.

Now let's make use of this neural

network and see how it works. Wonder how

traffic cameras identify vehicles

registration plate on the road to detect

speeding vehicles and those breaking the

law? They got me going through a red

light the other day. Well, last month.

That's like the horrible thing. They

send you this picture of you and all

your information because they pulled it

up off of your license plate and your

picture. I shouldn't have gone through

the red light. So, here we are and we

have an image of a car and you can see

the license plate on there. So, let's

consider the image of this vehicle and

find out what's on the number plate. The

picture itself is 28x 28 pixels and the

image is fed as an input to identify the

registration plate. Each neuron has a

number called activation that represents

the grayscale value of the corresponding

pixel range. And we range it from zero

to one. One for a white pixel and zero

for a black pixel. And you can see down

here we have an example where one of the

pixels is registered as like 082.

Meaning it's probably pretty dark. Each

neuron is lit up when its activation is

close to one. So as we get closer to

black on white, we can really start

seeing the details in there. And you can

see again the pixel shows us one up

there. It's like part of the car and so

it lights up. So pixels in the form of

arrays are fed to the input layer. And

so we see here the pixels of a car image

fed as an input. And you're going to see

that the input layer which is green is

one dimension while our image is

two-dimension. Now when we look at our

setup that we're programming in Python,

it has a cool feature that automatically

does the work for us. If you're working

with an older neural network pattern

package, you then convert each one of

those rows so it's all one array. So

you'd have like row one and then just

tack row two onto the end. You can

almost feed the image directly into some

of these neural networks. The key is

though is that if you're using a 28x 28

and you get a picture of this 30x30,

shrink the 30x30 down to fit the 28x 28.

So you can't increase the number of

input in this case green dots. It's very

important to remember when you work on

neural networks. And let's name the

inputs x1 x2 x3 respectively. So each

one of those represents one of the

pixels coming in. And the input layer

passes it to the hidden layer. And you

can see here we now have two hidden

layers in this image in the orange. And

each one of those pixels connects to

each one of those hidden layers. And the

interconnections are assigned weights at

random. So they get these random weights

that come through. If x1 lights up, then

it's going to be x1 times this weight

going into the hidden layer. And we sum

those weights. The weights are

multiplied with the input signal and a

bias is added to all of them. So as you

can see here we have X1 comes in and it

actually goes to all the different

hidden layer nodes or in this case uh

whatever you want to call them network

setup the orange dots and so you take

the value of X1 you multiply it by the

weight for the next hidden layer. So X1

goes to hidden layer 1 X1 goes to hidden

layer two X1 goes hidden layer 1 node

two hidden layer one node three and so

on. And the bias a lot of times they

just put the bias in as like another

green dot or another orange dot and they

give the bias a value one and then all

the weights go in from the bias into the

next node. So the bias can change. We

always just remember that you need to

have that bias in there. There's things

that can be done with it. Generally most

of packages out there control that for

you so you don't have to worry about

figuring out what the bias is. But if

you ever dive deep into neural networks,

you got to remember there's a bias or

the answer won't come out correctly. The

weighted sum of the input is fed as an

input to the activation function to

decide which nodes to fire. And for

feature extraction, as a signal flows

within the hidden layers, the weighted

sum of inputs is calculated and is fed

to the activation function in each layer

to decide which nodes to fire. So here's

our feature extraction of the number

plate. And you can see these are still

hidden nodes in the middle. And this

becomes important. We're going to take a

little detour here and look at the

activation function. So, we're going to

dive just a little bit into the math so

you can start to understand where some

of the games go on when you're playing

with neural networks in your

programming. So, let's look at the

different activation functions before we

move ahead. Here's our friendly red tag

shopping robot. And so, one is a sigmoid

function. And the sigmoid function which

is 1 over 1 + e to the minus x takes the

x value and you can see where it

generates almost a zero and almost a one

with a very small area in the middle

where it crosses over and we can use

that value to feed into another

function. So if it's really uncertain it

might have a 0.1 or 2 or 3 but for the

most part it's going to be really close

to one and really close to this case

zero zero to one the threshold function.

So if you don't want to worry about the

uncertainty in the middle, you just say,

"Oh, if x is greater than or equal to

zero, if not, then uh x is zero." So

it's either zero or one. Really

straightforward. There's no in between

in the middle. And then you have the

what they call the reel relu function.

And you can see here where it puts out

the value, but then it says, well, if

it's over one, it's going to be one. And

if it's uh less than zero, it's zero. So

it kind of just deadends it on those two

ends, but allows all the values in the

middle. And again, this like the sigmoid

function allows that information to go

to the next level. So it might be

important to know if it's a 0.1 or a

minus.1. The next hidden layer might

pick that up and say, "Oh, this piece of

information is uncertain or this value

has a very low certainty to it." And

then the hyperbolic tangent function.

And you can see here it's a 1 - e to the

-2x over 1 + e - 2x. And it's very much

along the same theme, a little bit

different in here in that it goes

between minus one and one. So you'll see

some of these it goes 0ero to one, but

this one goes minus one to one. And if

it's less than zero, it's, you know, it

doesn't fire and if it's over zero, it

fires. And it also still puts out a

value. So you still have a value you can

get off of that just like you can with

the sigmoid function and the relu

function. Very similar in use. And I

believe the originally used to be

everything was done in the sigmoid

function. That was the most uh commonly

used. And now they just kind of use more

the reloo function. The reason is one,

it processes faster because you already

have the value and you don't have to add

another compute the 1 / 1 + e to the

minus x for each hidden node and the

data coming off works pretty good as far

as putting it into the next level. If

you want to know just how close it is to

zero, how close is it not to

functioning, you know, is it minus.1

minus.2 usually they're float values.

You get like minus point minus.00138

or something. So, you know, important

information, but the Reu is most

commonly used these days as far as the

setup we're using. But you'll also see

the sigmoid function very commonly used

also. Now that you know what an

activation function is, let's get back

to the neural network. So, finally, the

model would predict the outcome of

applying a suitable activation function

to the output layer. So, we go in here,

we look at this, and we have the optical

character recognition OCR is used on the

images to convert it into a text in

order to identify what's written on the

plate. And as it comes out, you'll see

the red node. And the red node might

actually represent just the letter A. So

there's usually a lot of outputs when

you're doing text identification. We're

not going to show that on here, but you

might have it even in the order. It

might be what order the license plates

in. So you might have ABCDE E FG, you

know, all the alphabet plus the numbers.

And you might have the 1 2 3 4 5 6 7 8 9

10 places. So it's a very large array

that comes out. It's not a small amount

of uh, you know, we show three dots

coming in, eight hidden layer nodes, you

know, two sets of four. We just show one

red coming out. A lot of times this is

uh, you know, 28 * 28. If you did 30 *

30, that's, you know, 900 nodes. So 28

is a little bit less than that uh, just

on the input. And so you can imagine the

hidden layer is just as big. Each hidden

layer is just as big if not bigger. Then

the output is going to be there's so

many digits. You know, it's a lot.

There's it's a huge amount of input and

output. But we're only showing you just,

you know, it' be hard to show in one

picture. And so it comes up and this is

what it finally gets out in the output

as it identifies a number on the plate.

And in this case, we have 08-d3858.

Error in the output is back propagated

through the network and weights are

adjusted to minimize the error rate.

This is calculated by a cost function.

When we're training our data, this is

what's used and we'll look at that in

the code when we do the data training.

So, we have stuff we know the answer to

and then we put the information through

and it says yes, that was correct or no,

cuz remember we randomly set all the

weights to begin with. And if it's

wrong, we take that error. How far off

are you? You know, are you off by is it

if it was like minus one, you're just a

little bit off. If it's like minus 300

was your output, remember when we're

looking at those different options, you

know, hyperbolic or whatever, and we're

looking at the could doesn't have an

limit on top or bottom. it actually just

generates a number. So if it's way off,

you have to adjust those weights a lot.

But if it's pretty close, you might

adjust the weights just a little bit.

And you keep adjusting the weights until

they fit all the different training

models you put in. So you might have 500

training models and those weights will

adjust using the back propagation. It

sends the error backward. The output is

compared with the original result and

multiple iterations are done to get the

maximum accuracy. So, not only does it

look at each one, but it goes through it

and just keeps cycling through these the

data making small changes in the network

until it gets the right answers. With

every iteration, the weights at every

interconnection are adjusted based on

the error. We're not going to dive into

that math because it is a differential

equation and it gets a little

complicated, but I will talk a little

bit about some of the different options

they have when we look at the code. So,

we've explored a neural network. Let's

look at the different types of

artificial neural networks. And this is

like the biggest area growing is how

these all come together. Let's see the

different types of neural network. And

again, we're comparing this to human

learning. So here's a human brain. I

feel sorry for that poor guy. So we have

a feed for forward neural network.

Simplest form of a they call it a ann a

neural network. Data travels only in one

direction input to output. This is what

we just looked at. So as the data comes

in, all the weights are added, it goes

to the hidden layer, all the weights are

added, it goes to the next hidden layer,

all the weights are added, and it goes

to the output. The only time you use the

reverse propagation is to train it. So

when you actually use it, it's very

fast. When you're training it, it takes

a while because it has to iterate

through all your training data. And you

start getting into big data because you

can train these with a huge amount of

data. The more data you put in, the

better trained they get. The

applications vision and speech

recognition actually they're pretty much

everything we talked about a lot of

almost all of them use this form of

neural network at some level radio basis

function neural network this model

classifies a data point based on its

distance from a center point. What that

means is that you might not have

training data. So you want to group

things together and you create central

points and it looks for all the things

you know some of these things are just

like the other. If you've ever watched

the Sesame Street as a kid, that dates

me. So, it brings things together and

this is a great way if you don't have

the right training model, you can start

finding things that are connected you

might not have noticed before.

Applications power restoration systems.

They try to figure out what's connected

and then based on that they can fix the

problem if you have a huge power system.

and self-organizing neural network

vectors of random dimensions are input

to discrete map comprised of neurons. So

they basically find a way to draw they

call them they say dimensions or vectors

or planes because they actually chop the

data in one dimension, two dimension,

three dimension, four, five, six. They

keep adding dimensions and finding ways

to separate the data and connect

different data pieces together.

Applications used to recognize patterns

in data like in medical analysis. The

hidden layer saves its output to be used

for future prediction. Recurrent neural

networks. So the hidden layers remember

its output from last time and that

becomes part of its new input. Uh you

might use that especially in robotics or

flying a drone. You want to know what

your last change was and how fast it was

going to help predict what your next

change you need to make is to get to

where the drone wants to go.

Applications text to speech conversation

model. So, you know, I talked about

drones, but you know, just identifying

on Lexus or Google Assistant or any of

these, they're starting to add in I'd

like to play a song on my Pandora, and

I'd like it to be at volume 90%. So, you

now can add different things in there,

and it connects them together. The input

features are taken in batches like a

filter. This allows a network to

remember an image in parts. Convolution

neural network. today's world in photo

identification and taking apart photos

and trying to you know have you ever

seen that on Google where you have five

people together this is the kind of

thing separates all those people so then

it can do a face recognition on each

person applications used in signal and

image processing in this case I use

facial images or Google picture images

as one of the options modular neural

network it has a collection of different

neural networks working together to get

the output so wow we just went through

all these different types of neural

networks. And the final one is to put

multiple neural networks together. I

mentioned that a little bit when we

separated people in a larger photo and

individuals in the photo and then do the

facial recognition on each person. So

one network is used to separate them and

the next network is then used to figure

out who they are and do the facial

recognition. Applications still

undergoing research. This is a cutting

edge. you hear the term pipeline and

there's actual in Python code and in

almost all the different neural network

setups out there they now have a

pipeline feature usually and it just

means you take the data from one neural

network and maybe another neural network

or you put it into the next neural

network and then you take three or four

other neural networks and feed them into

another one. So how we connect the

neural networks is really just cutting

edge and it's so experimental. I mean

it's almost creative in its nature.

There's not really a science to it

because each specific domain has

different things it's looking at. So if

you're in the banking domain, it's going

to be different than the medical domain

than the automatic car domain. And

suddenly figuring out how those all fit

together is just a lot of fun and really

cool. So we have our types of artificial

neural network. We have our feed forward

neural network. We have a radial basis

function neural network. We have our

Cohen self-organizing neural network,

recurrent neural network, convolution

neural network, and modular neural

network where it brings them all

together. And u no the colors on the

brain do not match what your brain

actually does, but they do bring it out

that most of these were developed by

understanding how humans learn. And as

we understand more and more of how

humans learn, we can build something in

the computer industry to mimic that, to

reflect that. And that's how these were

developed. So exciting part, use case

problem statement. So this is where we

jump in. This is my favorite part. Let's

use the system to identify between a cat

and a dog. If you remember correctly, I

said we're going to do some Python code.

And you can see over here, my hair is

kind of sticking up over the computer,

cup of coffee on one side, and a little

bit of old school. A pencil and a pen on

the other side. Yeah, most people now

take notes. I love the stickies on the

computer. That's great. That's that is

my computer. I have sticky notes on my

computer in different colors. So, not

too far from uh today's programmer. So,

the problem is is we want to classify

photos of cats and dogs using a neural

network. And you can see over here we

have quite a variety of dogs in the

pictures and cats and you know just

sorting out it is a cat is pretty

amazing. And why would anybody want to

even know the difference between a cat

and a dog? Okay, you know why? Well, I

have a cat door. It'd be kind of fun

that instead of it identifying, instead

of having like a little collar with a

magnet on it, which is what my cat has,

the door would be able to see, oh,

that's the cat. That's our cat coming

in. Oh, that's the dog. We have a dog,

too. That's a dog I want to let in.

Maybe I don't want to let this other

animal in cuz it's a raccoon. So, you

can see where you could take this one

step further and actually apply this.

You could actually start a little

startup company idea, self-identifying

door. So, this use case will be

implemented on Python. I am actually in

Python 3.6. It's always nice to tell

people the version of Python because

that does affect sometimes which modules

you load and everything. And we're going

to start by importing the required

packages. I told you we're going to do

this in Kass. So we're going to import

from KAS models sequential from the Kass

layers conversion 2D or COV2D max

pooling 2D flatten and dense. And we'll

talk about what each one of these do in

just a second. But before we do that,

let's talk a little bit about the

environment we're going to work in. And

uh you know, in fact, let me go ahead

and open a uh the website, KASS's

website, so we can learn a little bit

more about KASS. So here we are on the

Kurass website, and it's uh ke.io.

That's the official website for Kurass.

And the first thing you'll notice is

that Kurass runs on top of either

TensorFlow, CNTK, and I think it's

pronounced Thano or Theo. What's

important on here is that TensorFlow and

the same is true for all these, but

TensorFlow is probably one of the most

widely used currently packages out there

with the KAS. And of course, you know,

tomorrow this is all going to change.

It's all going to disappear and they'll

have something new out there. So, make

sure when you're learning this code that

you understand what's going on and also

know the code. I mean, look, when you

look at the code, it's not as

complicated once you understand what's

going on. The code itself is pretty

straightforward. And the reason we like

KAS and the reason that people are

jumping on it right now, it's such a big

deal is if we come down here, let me

just scroll down a little bit. They talk

about user friendliness, modularity,

easy extensibility, work with Python.

Python's a big one because a lot of

people in data science now use Python,

although you can actually access Kass

other ways. Is if we continue down here

is layers. And this is where it gets

really cool. When we're working with

KASS, you just add layers on. Remember

those hidden layers we were talking

about? And we talked about the reelu

activation. You can see right here. Let

me just up that a little bit in size.

There we go. That's big. I can add in an

eelu layer. And then I can add in a

softmax layer in the next instance. We

didn't talk about softmax. So you can do

each layer separate. Now if I'm working

in some of the other kits I use, I take

that and I have one setup and then I

feed the output into the next one. This

one I can just add hidden layer after

hidden layer with the different

information in it which makes it very

powerful and very fast to spin up and

try different setups and see how they

work with the data you're working on.

And we'll dig a little bit deeper in

here. And a lot of this is very much the

same. So when we get to that part, I'll

point that out to you also. Now just a

quick side note, I'm using Anaconda with

Python in it. And I went ahead and

created my own package and I called it

the Kass Python 36 because I'm in Python

36. Anaconda is cool that You can create

different environments really easily. If

you're doing a lot of different

experimenting with these different

packages, probably want to create your

own environment in there. And the first

thing, as you can see right here,

there's a lot of dependencies. A lot of

these you should recognize by now if

you've done any of these videos. If not,

kudos for you for jumping in today. PIP,

install, numpy, sci, the scikitlearn,

pillow, and h5py

are both needed for the tensorflow and

then putting the kass on there. And then

you'll see here uh and pip is just a

standard installer that you use with

Python. You'll see here that we did pip

install TensorFlow since we're going to

do KAS on top of TensorFlow. And then

pip install and I went ahead and used

the GitHub. So git plusgit and you'll

see here github.com. This is one of

their releases, one of the most current

release on there that goes on top of

TensorFlow. And you can look up these

instructions pretty much anywhere. This

is for doing it on Anaconda. Certainly

you'd want to install these if you're

doing it in Iuntu server setup. you

you'd want to get I don't think you need

the H5 py and aru but you do need the

rest in there because they are

dependencies in there and it's pretty

straightforward and that's actually in

some of the instructions they have on

their website so you don't have to

necessarily go through this just

remember their website on there and then

when I'm under my uh Anaconda navigator

which I like you'll see where I have

environments and on the bottom I created

a new environment and I called it KAS

Python 36 just to separate everything

you can say I have Python 3.5 and Python

36 I used to have a bunch of other ones,

but it kind of cleaned house recently.

And of course, once I go in here, I can

launch my Jupyter Notebook, making sure

I'm using the right environment that I

just set up. This, of course, opens up

my um in this case, I'm using uh Google

Chrome. And in here, I could go and just

create a new document in here. And this

is all in your um browser window when

you use the Anaconda. Do you have to use

Anaconda and Jupyter Notebook? No. You

can use any kind of Python editor,

whatever setup you're comfortable with

and whatever you're doing in there. So,

let's go ahead and go in here and paste

the code in. And we're importing a

number of different settings in here. We

have import sequential. That's under the

models because that's the model we're

going to use as far as our neural

network. And then we have layers and we

have conversion 2D, max pooling 2D,

flatten dense. And you can actually just

kind of guess at what these do. We're

talking we're working in a 2D

photograph. And if you remember

correctly, I talked about how the actual

input layer is a single array. It's not

in two dimensions. It's one dimension.

All these do is these are tools to help

flatten the image. So, it takes a

two-dimensional image and then it

creates its own proper setup. You don't

have to worry about any of that. You

don't have to do anything special with

the photograph. You let the carass do

it. And we're going to run this. And

you'll see right here they have some

stuff that is going to be depreciated

and changed because that's what it does.

Everything's being changed as we go. You

don't have to worry about that too much.

If you have warnings, if you run it a

second time, the warning will disappear.

And this has just imported these

packages for us to use. Jupiter's nice

about this that you can do each thing

step by step. And I'll go ahead and also

zoom in there. A little control plus.

That's one of the nice things about

being in a browser environment. So, here

we are back. Another sip of coffee. If

you're familiar with my other videos,

you notice I'm always sipping coffee. I

always have a in my case latte next to

me, an espresso. So the next step is to

go ahead and initialize. We're going to

call it the CNN or classifier neural

network. And the reason we call it a

classifier is because it's going to

classify it between two things. It's

going to be cat or dog. So when you're

doing classification, you're picking

specific objects. You're specific. It's

a true or false. Yes, no. It is

something or it's not. So first thing

we're going to create our classifier and

it's going to equal sequential. So their

sequential setup is the classifier.

That's the actual model we're using.

That's the neural network. So we call it

a classifier. And uh the next step is to

add in our convolution. And let me just

do a uh let me shrink that down in size

so you can see the whole line. And let's

talk a little bit about what's going on

here. I have my classifier and I add

something. What am I adding? Well, I'm

adding my first layer. This first layer

we're adding in is probably the one that

takes the most work to make sure you

have it set correct. And the reason I

say that is this is your actual input.

And we're going to jump here to the part

that says input shape equals 64x 64x3.

What does that mean? Well, that means

that our pictures coming in. And there's

these pictures. Remember we had like the

picture of the car was 128x 128 pixels.

Well, this one is 64x 64 pixels. And

each pixel has three values. That's

where these numbers come from. And it is

so important that this matches. I

mentioned a little bit that if you have

like a larger picture, you have to

reformat it to fit this shape. If it

comes in as something larger, there's no

input notes. There's no input neural

network there that will handle that

extra space. So, you have to reshape

your data to fit in here. Now, the first

layer is the most important because

after that, KAS knows what your shape is

coming in here and it knows what's

coming out and so that really sets the

stage. Most important thing is that

input shape matches your data coming in.

And you'll get a lot of errors if it

doesn't. You'll go through there and

picture number 55 doesn't match it

correctly. And guess what it does? It

usually gives you an error. And then the

activation, if you remember, we talked

about the different activations on here.

We're using the reelu model. Like I

said, that is the most commonly used now

because one, it's fast. Doesn't have the

added calculations in it. It just says

here's the value coming out based on the

weights and the value going in. And um

from there, you know, it's uh if it's

over one, then it's good or over zero,

it's good. If it's under zero, then it's

considered not active. And then we have

this conversion 2D. What the heck is

conversion 2D? I'm not going to go into

too much detail in this because this has

a couple of things it's doing in here, a

little bit more in-depth than we're

ready to cover in this tutorial. But

this is used to convert from the photo

cuz we have 64x 64x3 and we're just

converting it to two-dimensional kind of

setup. So it's very aware that this is a

photograph and that different pieces are

next to each other. And then we're going

to add in uh a second convolutional

layer. That's what the cov stands for

2D. So it's these are hidden layers. So

we have our input layer and our two

hidden layers and they are

two-dimensional because we're dealing

with a two-dimensional photograph. And

you'll see down here that on the last

one, we add a max pooling 2D and we put

a pool size equals 22. And so what this

is is that as you get to the end of

these layers, one of the things you

always want to think of is what they

call mapping and then reducing.

Wonderful terminology from the big data.

We're mapping this data through all

these layers. And now we want to reduce

it to only two sets. In this case, it's

already in two sets because it's a 2D

photograph. But we had, you know, two

dimensions by we actually have 64x 64

by3. So now we're just getting it down

to a 2x two. Just the two dimension

two-dimensional instead of having the

third dimension of colors. And we'll go

ahead and run these. We're not really

seeing anything in our run script

because we're just setting up. This is

all set up. And this is where you start

playing because maybe you'll add a

different layer in here to do something

else to see how it works and see what

your output is. That's what makes KAS so

nice is I can with just a couple flips

of code put in a whole new layer that

does a whole new processing and see

whether that improves my run or makes it

worse. And finally, we're going to do

the final setup, which is to flatten

classifier, add a flatten setup. And

then we're going to also add a layer, a

dense layer, and then we're going to add

in another dense layer. And then we're

going to build it. We're going to

compile this whole thing together. So,

let's flip over and see what that looks

like. And we've even numbered them for

you. So, we're going to do the

flattening. And flatten is exactly what

it sounds like. We've been working in a

two-dimensional array of picture, which

actually is in three dimensions because

of the pixels. The pixels have a whole

another dimension to it of three

different values. And we've kind of

resized those down to 2x two. But now

we're just going to flatten it. I don't

want to have multiple dimensions being

worked on by tensor and by kas. I want

just a single array. So, it's flattened

out. And then step four, full

connection. So we add in our final two

layers. And you could actually do all

kinds of things with this. You could

actually leave out this some of these

layers and play with them. You do need

to flatten it. That's very important.

Then we want to use the dents again.

We're taking this and we're taking

whatever came into it. So once we take

all those different the two dimensions

or three dimensions as they are and we

flatten it to one dimension. We want to

take that and we're going to pull it

into units of 128. They got that. You're

say where did they get 128 from? You

could actually play with that number and

get all kinds of weird results. But in

this case we took the 64 + 64 is 128.

You could probably even do this with 64

or 32. Usually you want to keep it in

the same multiple whatever the data

shape you're already using is in. And

we're using the activation the re lu

just like we did before. And then we

finally filter all that into a single

output. And it has how many units? One.

Why? Because we want to know whether

true or false. It's either a dog or a

cat. You could say one is dog, zero is

cat. Or maybe you're a cat lover and

it's one is cat and zero is dog. And if

you love both dogs and cats, you're

going to have to choose. And then we use

the sigmoid activation. If you remember

from before, we had the reel and there's

also the sigmoid. The sigmoid just makes

it clear it's yes or no. We don't want a

any kind of in between number coming

out. And we'll go ahead and run this.

And you'll see it's still all in setup.

And then finally, we want to go ahead

and compile. And let's put the compiling

our um classifier neural network. And

we're going to use the optimizer atom.

And I hinted at this just a little bit

before. Where does atom come in? Where

does an optimizer come in? Well, the

optimizer is the reverse propagation.

When we're training it, it goes all the

way through and says error and then how

does it readjust those weights. There

are a number of them. Atom is the most

commonly used and it works best on large

data. Most people stick with the atom

because when they're testing on smaller

data, see if their model is going to go

through and get all their errors out

before they run it on larger data sets.

They're going to run it on atom anyway,

so they just leave it on atom most

commonly used. But there are some other

ones out there. You should be aware of

that that you might try them if you're

stuck in a bind or you might blur that

in the future, but usually atom is just

fine on there. And then you have two

more settings. You have loss and

metrics. We're not going to dig too much

into loss or metrics. These are things

you really have to explore KAS because

there are so many choices. This is how

it computes the error. There's so many

different ways to on your back

propagation and your training. So we're

using the atom model, but you can

compute the error by um standard

deviation, standard deviation squared.

They use binary cross entropy. I'd have

to look that up to even know what that

is. There's so many of these. A lot of

times you just start with the ones that

look correct that are most commonly used

and then you have to go read the KAS

site and actually see what these

different losses and metrics and what

different options they have. So, we're

not going to get too much into them

other than to reference you over to the

KAS website to explore them deeper, but

we are going to go ahead and run them.

And now we've set up our classifier. So,

we have an object classifier. And if you

go back up here, you'll see that we've

added in step one. We added in our layer

for the input. We added a layer that

comes in there and uses the reelu for

activation. And then it pulls the data.

So this is even though these are two

layers, the actual neural network layer

is up here. And then it uses this to

pull the data into a 2x two. So into a

two-dimensional array from a

three-dimensional array with the colors.

Then we flatten it. So there's our adder

flatten. And then we add another dense

what they call dense layer. this dense

layer goes in there and it it downsizes

it to 128. It reduces it. So you can

look at this as uh we're mapping all

this data down the two-dimensional setup

and then we flatten it. So we map it to

a flatten map and then we take it and

reduce it down to 128 and we use the

reel again. And then finally we reduce

that down to just a single output and we

use a sigmoid to do that to figure out

whether it's yes, no, true, false, in

this case cat or dog. And then finally

once we put all these layers together we

compile them. That's what we've done

here and we've compiled them as far as

how it trains to use these settings for

the training back propagation. So if you

remember we talked about training our

setup and when we go into this you'll

see that we have two data sets. We have

one called the training set and the

testing set. And that's very standard in

any data processing is you need to have

that's pretty common in any data

processing is you need to have a certain

amount of data to train it and then you

got to know whether it works or not. Is

it any good and that's why you have a

separate set of data for testing it

where you already know the answer but

you don't want to use that as part of

the training set. So in here we jump

into part two fitting the classifier

neuron network to the images and then

from KAS let me just zoom in there. I

always love that about working with

Jupyter Notebooks. You can really see.

We're going to come in here. We do the

cross pre-processing an image. And we

import image data generator. It's so

nice of KAS. It's such a high-end

product right now going out. And since

images are so common, they already have

all this stuff to help us process the

data, which is great. And so, we come in

here, we do train data gen, and we're

going to create our object for helping

us train for reshaping the data so that

it's going to work with our setup. and

we use an image data generator and we're

going to rescale it. And you'll see here

we have one point which tells us it's a

float value on the rescale over 255.

Where does 255 come from? Well, that's

the scale in the colors of the pictures

we're using. They're value from 0 to

255. So, we want to divide it by 255 and

it'll generate a number between 0 and 1.

They have sheer range and zoom range.

Horizontal flip equals true. And this,

of course, has to do with if the photos

are different shapes and sizes. Like I

said, it's a wonderful package. You

really need to dig in deep to see all

the different options you have for

setting up your images. For right now

though, we're going to just stick with

some basic stuff here. And let me go

ahead and run this code. And again, it

doesn't really do anything because we're

still setting up the pre-processing.

Let's take a look at this next set of

code. And this one is just huge. We're

creating the training set. So the

training set is going to go in here and

it's going to use our train data gen we

just created flow from directory. It's

going to access in this case the path

data set training set. That's a folder.

So it's going to pull all the images out

of that folder. Now I'm actually running

this in the folder that the data sets

in. So if you're doing the same setup

and you load your data in there and

you're doing this, make sure wherever

your Jupyter notebook is saving things

to that you create this path or you can

do the complete path if you need to, you

know, C colon slash etc. And the target

size, the batch size and class mode is

binary. So the classes, we're switching

everything to a binary value. Batch

size. What the heck is batch size? Well,

that's how many pictures we're going to

batch through the training each time.

And the target size 64x 64. A little

confusing, but you can see right here

that this is just a general training and

you can go in there and look at all the

different settings for your training

set. And of course with different data,

we're doing pictures. There's all kinds

of different settings depending on what

you're working with. Let's go ahead and

run that and see what happens. And

you'll see that it found 800 images

belonging to one classes. So we have 800

images in the training set. And if we're

going to do this with uh the training

set, we also have to format the pictures

in the test set. Now, we're not actually

doing any predictions. We're not

actually programming the model yet. All

we're doing is preparing the data. So,

we're going to prepare a training set

and the test set. So, any changes we

make to the training set at this point

also have to be made to the test set.

So, we've done this thing. We've done a

train data generator. We've done our

training set. And then we also have

remember our test set of data. So I'm

going to do the same thing with that.

I'm going to create a test data gen and

we're going to do this image data

generator. We're going to rescale one

over 255. We don't need the other

settings, just the single setting for

the test data gen. And we're going to

create our test set. We're going to do

the same thing we did with the test set

except that we're pulling it from the

test set folder. And we'll run that. And

you'll see in our test set we found

2,000 images. That's about right. We're

using 20% of the images as test and 80%

to train it. And then finally, we've set

up all our data. We've set up all our

layers, which is where all the work is

is cleaning up that data, making sure

it's going in there correctly. And we're

actually going to fit it. We're going to

train our data set. And let's see what

that looks like. And here we go. Let's

put the information in here. And let's

just take a quick look at what we're

looking at with our fit generator. We

have our classifier.fit

generator. That's our back propagation.

So the information goes through forward

with a picture and it says, "Oh, you're

either right or you're wrong." And then

the error goes backward and reprograms

all those weights. So we're training our

neural network. And of course, we're

using the training set. Remember, we

created the training set up here. And

then we're going steps per epic. So it's

8,000 steps. Epic means that that's how

many times we go through all the

pictures. So we're going to rerun each

of the pictures. and we're going to go

through the whole data set 25 times, but

we're going to look at each picture

during each epic 8,000 times. So, we're

really programming the heck out of this

and going back over it. And then they

have validation data equals test set.

So, we have our training set and then

we're going to have our test set to

validate it. So, we're going to do this

all in one shot and we're going to look

at that and they're going to do 200

steps for each validation and we'll see

what that looks like in just a minute.

Let's go ahead and run our training

here. And we're going to fit our data.

And as it goes, it says epic one of 25.

You start realizing that this is going

to take a while. On my older computer,

it takes about 45 minutes. I have a dual

processor. You know, we're processing uh

10,000 photos. That's not a small amount

of photographs to process. So, if you're

on your laptop, you know, which I am,

it's going to take a while. So, let's go

ahead and uh go get our cup of coffee

and a sip and come back and see what

this looks like. So, I'm back. You

didn't know I was gone. That was

actually a lengthy pause there. I made a

couple changes. Let's discuss those

changes real quick and why I made them.

So, the first thing I'm going to do is

I'm going to go up here and insert a

cell above and let's paste the original

code back in there. And you'll see that

the original thing was steps per epic

8,000, 25 epics, and validation steps

2,000. And I changed these to 4,000

epics or 4,000 steps per epic, 10 epics,

and just 10 validation steps. And this

will cause problems if you're doing this

as a commercial release. But for demo

purposes, this should work. And if you

remember our steps per epic, that's how

many photos we're going to process. In

fact, let me go ahead and get my drawing

pen out. And uh let's just highlight

that right here. We have 8,000 pictures

we're going through. So for each epic,

I'm going to change this to 4,000. I'm

going to cut that in half. So, it's

going to randomly pick 4,000 pictures

each time it goes through an epic. And

the epic is how many processes. So, this

is 25. And I'm just going to cut that to

10. So, instead of doing 25 runs through

8,000 photos each, which you can do the

math of 25 * 8,000, I'm only going to do

10 through 4,000. So, I'm going to run

this 40,000 times through the processes.

And the next thing I not you'll you'll

want to notice is that I also changed

the validation step. And this would

cause some major problems in releasing

cuz I dropped it all the way down to 10.

What the validation step does is it says

we have 2,000 photos in our training or

in our testing set and we're going to

use that for validation. Well, I'm only

going to use a random 10 of those to

validate. So, not really the best

settings, but let me show you why we did

that. Let's scroll down here just a

little bit and let's look at the output

here and see what that what's going on

there. So, I've got my drawing tool back

on, and you'll see here it lists a run.

So, each time it goes through an epic,

it's going to do 4,000 steps. And this

is where the 4,000 comes in. So, that's

where we have. We have epic one of 10,

4,000 steps. So, it's randomly picking

half the pictures in the file and going

through them. And then we're going to

look at this number right here. That is

for the whole epic, and that's 24, 411

seconds. And if you remember correctly,

you divide that by 60, you get minutes.

If you divide that by 60, you get hours.

Or you can just divide the whole thing

by 60 * 60 which is 3600. If 3600 is an

hour, this is roughly 45 minutes right

here. And that's 45 minutes to process

half the pictures. So if I was doing all

the pictures, we're talking an hour and

a half per epic times 36 or no 25. They

had 25 up above 25. So that's roughly a

couple days. A couple days of

processing. Well, for this demo, we

don't want to do that. I don't want to

come back the next day. Plus, my

computer did a reboot in the middle of

the night. So, we look at this and we

say, "Okay, let's we're just testing

this out. My computer that I'm running

this on is a dual core processor. Uh,

runs 0.9 gigahertz per second. For a

laptop, you know, it's good about 4

years ago, but for running something

like this, it's probably a little slow.

So, we cut the times down. And the last

one was validation. We're only

validating it on a random 10 photos. And

this comes into effect because you're

going to see down here where we have

accuracy, value loss, value accuracy,

and loss. Those are very important

numbers to look at. So the 10 means I'm

only validating across 10 pictures. That

is where here we have value. This is ACC

is for accuracy. Value loss. We're not

going to worry about that too much. And

accuracy. Now accuracy is while it's

running, it's putting these two numbers

together. That's what accuracy is. And

value accuracy is at the end of the

epic. What's our accuracy into the epic?

What is it looking at? In this tutorial,

we're not going to go so deep, but these

numbers are really important when you

start talking about these two numbers

reflect bias. That is really important.

We just put that up there. And bias is a

little bit beyond this tutorial, but the

short of it is is if this accuracy,

which is being our validation per step

is going down and the value accuracy

continues to go up, that means there's a

bias. That means I'm memorizing the

photos I'm looking at. I'm not actually

looking for what makes a dog a dog, what

makes a cat a cat. I'm just memorizing

them. And so the more this discrepancy

grows, the bigger the bias is. And that

is really the beauty of the KAS neural

network. It has a lot of built-in

features like this that make that really

easy to track. So let's go ahead and

take a look at the next set of code. So

here we are into part three. We're going

to make a new prediction. And so we're

going to bring in a couple tools for

that. And then we have to process the

image coming in and find out whether

it's an actual dog or cat if we can

actually use this to identify it. And of

course the final step of part three is

to print prediction. We'll go ahead and

combine these. And of course you can see

me there adding more sticky notes to my

computer screen hidden behind the

screen. And you know last one was don't

forget to feed the cat and the dog.

So let's go and take a look at that and

see what that looks like in code and put

that in our Jupyter notebook. All right.

And let's paste that in here. And we'll

start by importing numpy as np. Numpy is

a very common package. I pretty much

import it on any Python project I'm

working on. Another one I use regularly

is pandas. They're just ways of

organizing the data. And then np is

usually the standard in most machine

learning tools as the return for the

data array. Although you know you use a

standard data array from Python. And we

have cross pre-processing import image.

This should all look familiar because

we're going to take a test image and

we're going to set that equal to in this

case cat or dog one as you can see over

here. And you know let me get my drawing

tool back on. So let's take a look at

this. We have our test image we're

loading and in here we have test image

one. And this one hasn't data hasn't

seen this one at all. So this is all

new. Oh, let me shrink the screen down.

Let me start that over. So here we have

my test image and we went ahead and the

cross processing has this nice image

setup. So we're going to load the image

and we're going to alter it to a 64x 64

print. So right off the bat, we're going

to cross is nice that way. It

automatically sets it up for us so we

don't have to redo all our images and

find a way to reset those. And then we

use also to set the image to an array.

So again, we're all in pre-processing

the data just like we pre-processed

before with our test information and our

training data. And then we use the

numpy. Here's our numpy that's uh from

our um right up here. Import numpy as in

p expand the dimensions test image axis

equal zero. So it puts it into a single

array. And then finally all that work

all that pre-processing and all we do is

we run the result. We click on here we

go result equals classifier predict test

image. And then we find out, well, what

is the test image? And let's just take a

quick look and just see what that is.

And you can see when I ran it, it comes

up dog. And if we look at those images,

there it is. Cat or dog. Image number

one. That looks like a nice floppy eared

lab. Friendly with his tongue hanging

out. It's either that or a very floppy

eared cat. I'm not sure which. But

according to our software, it says it's

a dog. And uh we have a second picture

over here. Let's just see what happens

when we run the second picture. We can

go up here and change this uh from dog

image one to two. We'll run that. and it

comes down here and says cat. You can

see me highlighting it down there as

cat. So, our process works. You're able

to label a dog a dog and a cat a cat

just from the pictures. There we go.

Cleared my drawing tool. And the last

thing I want you to notice when we come

back up here to when I ran it, you'll

see it has an accuracy of one and the

value accuracy of one. Well, the value

accuracy is the important one because

the value accuracy is what it actually

runs on the test data. Remember, I'm

only testing it on. and I'm only

validating it on a random 10 photos and

those 10 photos just happened to come up

one. Now, when they ran this on the

server, it actually came up about 86%.

This is why cutting these numbers down

so far for a commercial release is bad.

So, you want to make sure you're a

little careful of that when you're

testing your stuff that you change these

numbers back when you run it on a more

enterprise computer other than your old

laptop that you're just practicing on or

messing with. And we come down here and

again, you know, we had the validation

of cat. And so we have successfully

built a neural network that could

distinguish between photos of a cat and

a dog. Imagine all the other things you

could distinguish. Imagine all the

different industries you could dive into

with that. Just being able to understand

those two difference of pictures. What

about mosquitoes? Could you find the

mosquitoes that bite versus the

mosquitoes that are friendly? It turns

out the mosquitoes that bite us are only

4% of the mosquito population, if even

that, maybe 2%. There's all kinds of

industries that use this and there's so

many industries that are just now

realizing how powerful these tools are.

Just in the photos alone, there is a

myriad of industries sprouting up. And I

said it before, I'll say it again. What

an exciting time to live in with these

tools and that we get to play with. So

key takeaways. Well, we covered what is

a neural network. We use all kinds of

processing the map images on your phone.

We talked about things that a neural

network can do. translate text, identify

faces all the way to control robots, you

know, lots of exciting things. How does

a neural network work? So, we discussed

that with the different layers going

from the picture to the input layer to

the hidden layers and their weights to

the final output layer. We also talked

about how it does the math and computing

the output as yes or no, categorically

true false. We discussed types of

artificial neural networks. A lot of

vocabulary there from the feed forward

neural network which is the most

commonly used. That's the one the neural

network we used is a feed forward neural

network that does backward propagation

to train. And there's a lot of other

ones out there. There's the radial

biases, the cohen self-organizing

recurrent neural network, convolution

neural network, modular neural network.

The big one was modular because it

incorporates pieces of all the other

ones. So that whatever you're working on

now is a huge conglomerate of multiple

networks. Just all cutting edge. All of

it's new. People even working on it

don't even know where it's going. Again,

very exciting times. And finally, we dug

through my favorite part. You can see

with my uh latte on one side, my old

school pens and pencil, and all my

sticky notes working away. That's not

actually me, by the way. You probably

guessed that. And we walked through and

actually did a cat and dog photo, a

simple cat and dog photo. And you could

see where some of the problems are in

processing large amounts of photographs

and data where that starts to become

going from a single machine on my laptop

with this, you know, lower amount of

resources all the way to big data. How

if you're processing hundreds and

thousands of these photos, this now

needs to be set up on an enterprise

machine or even on a cluster of

computers. Again, significantly past the

scope of this. The neat part about it

though is once you write this code, most

of this code, they now have tools that

you can almost take the same ideas, if

not the actual code, and push it right

onto a cluster computation. So really

cool times for this Python. My name is

Richard Kersner with the SimplyLearn

team. That's www.simplearn.com.

Get certified, get ahead. Although deep

learning is uh been around for a while,

it is just in its infant stages of

development as far as exploding on the

market. I mean it is right now they're

building robots with it. Deep learning

is used to train robots to perform human

tasks. Music composition. Deep neural

nets can be used to produce music by

making computers learn the patterns

involved in composing music. Image

colorization. Neural network recognizes

objects and uses information from the

images to color them. Machine

translation. Given a word, phrase or a

sentence in one language, neural

networks automatically translate them

into another language. Google Translate

is one such popular machine translator

you may have come across. And you'll

notice in here we didn't show any

examples of straight numbers like uh

projective cells in a business tracking

your favorite stock. You can certainly

do those with machine languages, but

this is the next level. Uh save that for

your regression models, your linear

regression where you're actually

processing and crunching just straight

numbers. With machine learning and deep

learning, we're going to a whole new

level as far as what we can figure out

on the computer. What's in it for you?

We're going to cover what is deep

learning. We're going to take a look at

the biological versus artificial

intelligence. What is neural network

activation function in your neural

network and the cost function and how do

neural networks work. How do neural

networks learn? So there's a little

you'll see a switch right there. We just

went from how are they working in the

math in the background to exactly how

are they learning. We'll be implementing

the neural network. We'll do a gradient

descent deep learning platforms and

we'll give an introduction to TensorFlow

and implementation in TensorFlow. That's

Google's platform that they open sourced

recently and it's probably one of the

most cutting edges in deep learning and

even it is still in the infant stage

which is one of the reasons they

released it to open source. What is deep

learning? Deep learning is a sub field

of machine learning that deals with

algorithms inspired by the structure and

function of the brain. And you can see

we have a nice picture here. We have

artificial intelligence which is kind of

the big bubble that encompasses all

these different things we're talking

about. This is ability of machine to

imitate intelligent human behavior. And

in there we have machine learning

application of AI that allows a system

to automatically learn and improve from

experience. And if you looked at any of

our other videos, you'll know that

machine learning covers a lot. So deep

learning is a subcategory of that. But

don't forget machine learning has all

kinds of other tools that people use to

do very basic uh descriptive and

predictive and postcriptive uh

analytics. And then you have deep

learning application of machine learning

that uses complex algorithms and deep

neural nets to train a model. Let's take

a look at the biological neuron versus

the artificial neuron. Now remember in

the human brain and and this is true for

most animals there are a lot of

different neurons going on. So this is

the very basic one. I mean there's

hundreds of different cells involved. So

when we talk about neural networks and

this is why I say it's in a very infant

stage. They're really basing it on uh

just the most basic thing that we're

able to figure out going on in the

neural networks. And you can see right

here we have dendrites fetch information

from an adjacent neurons and pass them

on as inputs. So you have your data

coming in and your data going out. Any

computer model should be looking at that

what's coming in what's going out. The

data is fed as an input to the neuron.

So we look at the artificial neuron. You

can see we have our inputs. They come in

each one is specially weighted into the

neuron and then the neuron has an

output. The cell nucleus processes the

information received from the dendrites

and the neuron processes the information

provided as inputs. Axons are the cables

over which the information is

transmitted and the information is

transferred over weighted channels. So

you can look at that uh I mentioned

weights briefly but you alter the data

coming in. So those weights are what

causes different information coming in

to be weighted differently and processed

differently. And the synapses receive

the information from the axons and

transmit it to the adjacent neurons.

That's in your biological model. And

then when we look at the artificial

neuron, the output is a final value

predicted by the artificial neuron. So

as we dig deeper into looking at the

theory behind the neural network and we

kind of flip back and forth between

these because there's two huge aspects

of it. One is from the outside. What are

you seeing and what's going on from the

inside so you can find to do what you

need to do and give the best results you

can. And we start off with what do we

feed? We feed an unlabeled image to a

machine which identifies it without any

human intervention. And so you can see

here we have a circle that comes in at

784 pixels and it comes in by 28x 28.

And you can see how it colors in the um

the circle on there. And we put a

triangle in. The triangle in also comes

in as 28x 28 and it has 784 pixels. So

you'll see between these two both of

them are 784 pixels. This machine is

intelligent enough to differentiate

between the various shapes. So that's

what we want to use our neural network

to do is to say hey this is a circle.

This is a triangle. That's more of a

categorical. You can also do a

regression model where you're actually

putting out float value or a numerical

value. We'll be looking at the true

false or the categorical model mostly

because that's where you usually start

at the different there is no real

difference when you as far as the way

the internal functioning goes when you

start flipping between them other than

well we'll talk about that in just a

minute. So you can actually go between

the two quite easily and the neural

network provides this capability. So

we're going to use this capability to

look between those two. One of the

things I want you to note in here is

that we're looking at 784 pixels. We're

looking at 784 inputs. That's very

different than stock with a high low or

last year's sales based on date or we're

looking at just a couple of numbers and

they're very clear. They're numbers.

They're very clear what they are, which

is something you'd put into a machine

learning linear regression model. This

is a step up from that in that we're

looking at complex patterns and how do

you figure those complex patterns out.

So, a neural network is a system modeled

on the human brain. And we looked at

that comparing the two. Let's go ahead

and look deeper into the neural network

itself. We have our inputs coming in. So

the inputs are fed to a neuron that

processes a data and gives us an output.

Input and output. This is the most basic

structure of a neural network known as a

perceptron. So if you see the term

perceptron, that's what we're talking

about. We're talking about this single

node that has inputs and an output.

However, neural networks are usually

much more complex. Let's start with

visualizing a neural network as a black

box. And I always love that symbol. It's

a black box. It's kind of magical. We

have our inputs coming in and we want

certain outputs. The box takes inputs,

processes them, and gives an output.

Let's have a look at what happens within

this box. And you can see me there in my

uh secret agent getup and I got my

hidden hood and everything. I guess I'm

part of the uh black skull or something

like that group. Uh so let's take a look

at what happens within this magic box.

And remember, we're skipping back and

forth between the theory of what's going

on in the box, which you have to know

how to fine-tune and how to build,

versus looking at it from the outside.

We're programming this box, and we have

an input and an output to the box as a

whole. Within the box exists a network

that is a core of deep learning. And you

can see here we're showing one layer and

we have our grid coming in. The network

consists of layers of neurons. Each

neuron is associated with a number

called the bias. And you can think of

the bias uh if you overly simplify this

and we're doing a linear regression

model. This is your y intercept in your

uklidian geometry. You have to have

something that offsets it. And so you

always have a bias in these cells.

Neurons of each layer transmit

information to neurons of the next layer

over channels. And so you can see each

of our layers going through from left to

right. These channels are associated

with numbers called weights. These

weights along with the biases determine

the information that is passed over from

the neuron to neuron. So just like the

bias is your y intercept in uklidian

geometry. You could look at the an one

weight. Remember this is very

complicated. So we're not looking at

just one weight. You could look at the

weight as your slope of the line. Or if

you're doing x= uh my y + c, it would be

the m value. Neurons of each layer

transmit information to neurons of the

next layer. And you can see here as they

light up going across into the final

layer. and then to the output. And in

this case, the output is going to be

either uh a square in this one or it

might light up the other one which is a

circle. The output layer emits a

predicted output. So in this case, we're

looking at a classification uh true

false. Is it a circle? Is it a triangle?

Is it a square? Let's now go deeper.

What happens within the neuron? So we're

going to dig deeper and start getting a

little bit closer to some of the math.

Don't worry, you don't have to be a

calculus expert and know your

differential equations. Even though this

is one giant differential equation, you

don't need to understand those to

understand what's going on. Within each

neuron, the following operations are

performed. The product of each input and

the weight of the channel it's passed

over is found. This is simply addition.

We're going to sum up the weight times

the output from the previous channel and

plus the bias. Sum of the weighted

products is computed. This is called the

weighted sum. Bias unique to the neuron

is added to the weighted sum. The final

sum is then subjected to the particular

function and we'll discuss those that

particular function. That part is really

important because those functions uh

have a huge impact on how well your

model performs under different

conditions. The final sum is then

subject to a particular function. This

is the activation function. So if you

ever hear the term activation function,

that's what we're talking about. What

activates this cell and what doesn't. As

we dig deeper into activation function,

an activation function takes the

weighted sum of the input as its input

adds a bias and provides an output. And

a lot of times you'll actually see one

formula for the sum of the weight the

weighted sum and the bias. You'll just

see that as a single line of everything

added together. And here we've broken it

apart because it makes it clear that

this bias is not computed the same as

the weighted sums. Here are the most

popular types of activation function.

And I always find these interesting

because at one point I was sitting at a

table with a gentleman who was finishing

his PhD. He was in his last year and he

said he went through all this stuff and

he ended up just trying the four

different activation functions on this

particular problem he was working on. So

knowing the math behind it doesn't

necessarily mean you're going to know it

right away. Uh so even somebody who

might have a PhD and be doing the

calculations on this comes back out of

it and ends up just trying the different

uh um activation functions to see what's

going to make a difference. And a lot of

times that's a final step. That's the

kind of thing where you built your whole

model. You've come back and you're like

wait a minute can I do a better deal

with a sigmoid function or the threshold

or the rectifier. Knowing what they're

doing is important so you can explain it

to somebody else. And again you probably

do this on a small set of data. If

you're working with big data, uh you

don't want to take down the full server

farm just to test out your three

different series. You take a small

portion of that data, test it, and then

you put it through to the big data. So

let's take a look at this. We have the

sigmoid function, and it's used for

models where we have to predict the

probability as an output. It exists

between zero and one. And you'll see

that's true of all of our activation

functions we're working with. Either the

cells on or off, it's true or false. And

there might be a little variation in

there which as an output could be used

to compute uncertainty in your solution.

So if you're getting a 7 with this

activation function, it might be well

I'm not sure if that's really a square

or I'm not sure that's really a

triangle. And that might be a flag for

it to be looked at by a human observer

at least in today's models where we're

at right now. And you can see here we

have the formula is simply equals 1 over

1 + e the minus x where x is your value

coming in. and it's going to give you a

result that looks very similar to the

graph on there which is somewhere

between zero and one. Um, and right in

the middle you can see that there's a

huge uh kind of you can go through all

the different values and uncertainties

involved. So the sigmoid function is

probably the default on most of them. Uh

the next one is the threshold function.

It is a threshold-based activation

function. If x value is greater than a

certain value, the function is activated

and fired. Else not. Pretty

straightforward. Yes, no, true, false.

um I don't want to test for

improbabilities. I just want a straight

answer. I don't want to know if there's

a partial value on there. It either is

true or it's false. And the rectifier

function, it is the most widely used

activation function. I would debate

that. Um rectifier is pretty common one,

although I see that the sigmoid function

is used to be the basic one, but it's up

there. The rectifier function is very

commonly used. You get the output of X

if X is positive and zero otherwise. And

you can see here again just like um uh

it's either you kind of get a value

going up there. So max of x of zero. So

it's it's again it's like the threshold

function. Yes, no, true, false. Uh it's

either zero or it's uh some kind of

progressive value. And then we have the

rectifier function. I would argue with

this because the sigmoid function used

to be the most common one. But with the

rectifier function, it now says it is

the most commonly used or widely used

activation function and gives an output

of X if X is positive and zero

otherwise. This is kind of nice because

it now says absolutely not or it gives

you a value of probability. Now, when I

say a value of probability, be very

careful there. I'm not saying that it's

going to tell you this is 75% chance of

being a circle. I'm going to tell you

that it says, hey, if this says 0.1, it

probably needs to be looked at or 2 or

3. It's going to depend on your data as

to what that value means. In general,

that just means it's flagging it that if

it's not a one, then chances are it

needs to be looked at by a person and

re-evaluated. And there's a hyperbolic

tangent function. This function is

similar to sigmoid function is bound to

a range of minus1 to 1. So you can see

there's our 1 - eus 2x and 1 plus over 1

+ eus 2x. Again, it's very similar to

the sigmoid function. The bonus of the

hyperbolic function is you have that

variable coming through the middle. So

again, you can look at it and you have a

little bit more weight as far as you can

process that down the line. That's a

little bit more advanced than than what

we're looking at right now. And a lot of

times it's not even necessary in a lot

of our different uh uses for these

activation functions. Now, we looked at

activation functions and I kind of said

those are a little bit like a black box

because even if you know all the math, a

lot of times you end up just playing

with them to find out what works. And it

also depends on what model you're

working with, whether you need a flat

yes, no, true, false, or you need to

have something in the middle that says,

hey, this isn't quite a one. You might

need to process this with the human

intervention. And you could look at

that. Uh, one example would be

self-driving cars. You don't want a car

to be yes, no, I'm going to go through

the the light. You want it to be like,

okay, if it's uh almost yes, maybe we

stop and have human intervention so we

don't get an accident. Cost function is

something you can really see and measure

and is very important. The cost value is

the difference between the neural net's

predicted output and the actual output

from a set of labeled training data. So

we have our group of data that's a

square circle and since we're looking at

geometrical shapes, we've had somebody

already labeled that data. They've

already said this is a triangle, this is

a square. And so if this is coming up

and it's giving us and it's saying a

square is a triangle and it's saying a

triangle is a circle, the output is

wrong. And so that output can then be

measured in the versus the actual output

and that's the cost. Uh you might also

hear this as error because that's the

error value being returned. How far off

is it? And what we're looking for is the

least cost or the least error value. And

it's obtained by making adjustments to

the weights and biases iteratively

throughout the training process. And

this this is called back propagation.

And we're going to look in that a little

deeper as we look into an example. It's

really hard to see when you're just

looking at arrows without actual numbers

and where that flow is coming from. But

you can look at this is here's our

inputs. They put out a prediction. The

prediction comes out and says, "Hey,

we've already labeled this data cuz

we're in training mode and the training

data is off. This is the cost. Can we

send that error or that cost back and

adjust those weights?" And we do it in

very small increments across large

amounts of data so that those weights

minimize that cost or that error. But

what happens within these neurons? So

let's look at a little example of this.

Kind of helps if you have some kind of

visual. Let's build a neural network

that predict bike prices based on a few

of its features. And we'll see here we

have our CC, our mileage, and our ABS.

And these are our three input layers.

And then we have the bike price and the

output layer. Now, it doesn't do us very

good to just uh pump it in from the

beginning and pump it out. And to be

honest, I would use a machine learning

linear regression model on this since

these are just straight numbers. But

because we want a simple example, we're

going to put this through and show you

as a neural network what that looks

like. And we got to put a hidden layer

in there. The hidden layer helps in

improving the output accuracy. And you

could look at this as a bunch of ores.

So it might say, hey, when we compare

these three values on the first hidden

layer neuron, we're looking at one set

of features and then we might weight

them in the second one. So these are a

bunch of different ores kind of how the

math comes out in behind the scenes. And

then they go out of course to the bike

trace or the output layer. And each of

the connections have a weight assigned

with it. And you'll see here we have a

mileage CC with the weight one and

weight two going into our first neuron.

And you'd also have your ABS going in

there. And so X1 * weight 1 + X2 *

weight 2 plus the bias of one. And step

two is our activation. The activation

function coming in there. When does this

fire? And the neuron takes a subset of

the inputs and processes it. And then we

go through and we do that with the um

second hidden layer neuron and the third

one and so on. So you process each layer

in order going forward. Now when I told

you this is in its infant stage, they

now have neurons that fire into the same

layer or back a layer so that you now

have a time series and there's all kinds

of wild things that they're

experimenting with on these layers. This

basic setup has been around since the

mid90s. It's only now because of our

technology that it's open to almost

everybody to play with it. And that's

why I say this is in an infant stage in

development is this basic math is here,

but what we can do with it is amazing.

And what they're actually doing with all

these different things is amazing. And

so we're just at the beginning of how to

use all these different tools and our

deep learning and our neural networks.

Uh and so once we have our hidden layer

computed, the information reaching the

neurons in the hidden layer is subjected

to the respective activation function.

And so each one of these fires an

activation output uh and then those are

each weighted to the final output layer.

So the processed information is now sent

to the output layer once again over

weighted channels. And you could look at

this as each one of these is um I always

look at this as like a group of people.

They're all looking at the bulletin

board and the first person says this is

what I project sales for the company and

the second person and the third and so

on. And then their perspectives are

weighted based on their expertise. So

your accountant might have a very high

weight where the um maybe your janitor

has a very low weight because their

expertise is not in accounting and then

that goes into the output layer and once

in the output layer it goes uh the

output which is the predicted value is

compared against the original value. So

now we have our output layer and since

we have like already a list of uh bikes

with their the different setups and what

their value is we can now generate an

error from this. The cost function

determines the error in prediction and

reports it back to the neural network.

So this is the cost. This is how far off

it is. This is your error coming back.

And as you can see, this is back

propagation going on. So now our error

is going in reverse because we know

we're not completely correct on this

particular channel. The weights are

adjusted in order to reduce the error.

So each time we go back, we are changing

those weights to reduce that error. and

we change them in small increments. You

don't want to fit one input. Remember,

you might have a data pool with a

terabyte of data. You don't want to

solve for the first set of data that

comes in and that be the main solution

because everything else will be off.

This is going to confuse you. That's

also called a bias. So, we have the bias

in the cell where we're adding a value,

the kind of like the y intercept, and we

have a bias of the whole neural network,

which means that it's weighted towards

one set of answers. So we want to make

small changes in these weights so we

don't create a bias and the weights are

adjusted in order to reduce the error or

the cost. The network is now trained

using the new weights. Once again the

cost is determined and back propagation

is continued until the cost cannot be

reduced any further. So let's go ahead

and plug in values and see how our

neural network works. So here we come in

here and initially our channels are

assigned with random weights. This is

important because if you assign them all

with the same weight, you might be able

to reproduce it. But it turns out that

if I put all my weights as one or all my

weights as zero, it takes longer to

train where if you have random weights,

they already have like a little bit of

adjustment and ores built in and that

will give us a better answer and train

faster. Our first neuron takes a value

of mileage and CC as inputs. So here

comes our computation whatever those

inputs are. And we do that again with

the second neuron with those values

coming in. You can see here we have

weight three and so on and then our

third neuron coming down and of course

our fourth neuron. So we're adding all

these different values coming in here in

our hidden layer. The process value from

each neuron is sent to the output layer

over weighted channels. So again here's

our weights coming in and we have N1,

N2, N3 and N4. Once again the values are

subjected to the activation function and

a single value is emitted as the output.

On comparing the predicted value to the

actual value, we clearly see that our

network requires training. So, here we

have it that our bike price uh we put

out, we thought it was worth 2,000 on

our random weights and the bike actually

was $4,000 on there. Guessing that's not

US dollars cuz that'd be a very

expensive bike. But maybe it is. There's

some $2,000 $4,000 bikes out there. The

cost function is calculated and back

propagation takes place. And this is

pretty simple. You can look at that as

our um we're subtracting one value from

the other. We square it and then we take

half of that and that is propagated back

up. And each layer generates its own

errors. Let's go back one because you

have your predicted Y and your actual Y.

That goes back to the first layer. And

then based on the value of the cost

function, certain weights are changed.

So when we look at the next layer, that

error is not the original 4,000 - 2,000

squar / 2. This error is based on the

error of each cell generated. How far

off is that cell as far as its weights.

We're not going to show you. It's

actually a very complicated differential

equation. And you can probably write it

out if you wanted to. You just write out

each formula that goes into the next

level and you add them all together and

you can write it out all the way

through. Computers make it so you don't

have to. And our neural network is

considered trained when the value for

the cost function is minimum. So when we

get our error way down as low as we can,

that's when our neural network is

trained. And there I mean just recently

they've come up with all kinds of

different means for measuring that

particular value. a little bit beyond

the scope of today's neural network, but

you can actually you can actually see

it. You know, how far do you do this

until the neural network doesn't need to

be trained anymore and you can overtrain

a neural network. Now, the tools that

we're looking at automatically let you

know when to stop, which is really nice.

And that is just like I said, we're at

the beginning stages in neural networks

and it's just really cool what they can

do now and how much of it's automated

and how much of it is experimental.

Right now, let's take a look at gradient

descent. But what approach do we take to

minimize the cost function? So here we

have nice error thing coming in. This is

our cost or our error. Uh let's start

with plotting the cost function against

the predicted value. And so you can see

they fed in multiple y's and these are

the errors coming in and the cost of

each of these inputs and changes going

on. Note we start at a random point on

the curve. So usually you put in you

know you pick up your data and you

randomly pick where to start in your

data. A lot of times you just run it

from the beginning because you're going

through so much data, it's not that big

of a deal. But you start with one point

going in. So your forward propagation

goes through. You're going to go ahead

and find your cost or your error. It

points that on the curve. And you can

see how we're plotting it right here.

Since the gradient at this point is

positive, we may move right. So we're

going to move a little bit to the right

on here. And this time the gradient is

negative. We move a little bit to the

left. Eventually we try out the point

where the gradient is zero. This is a

least value of cost function. You have

to be a little careful with this because

this particular I mean they make it look

nice and simple in this graph. Sometimes

these curves look like stair steps and

so there is global minimums and then

there is local there might be a local

point where the gradient is zero but

it's not the global one. Uh so it might

be way off to the left where it just

happens to step down a little bit and

you think you're in the right gradient.

And with that we have all the right

weights and we can say our network is

trained. So here we have um just some

major these are some of the big names

out there right now in development for

deep learning platforms. TensorFlow

which we'll actually do an example in in

a minute. Deep learning for J which is

in the Java platform. Uh so if you're a

Java programmer uh by the way is

TensorFlow is accessed most people are

using Python to access it but it is a

system that's kind of separate from a

lot of the programming languages which

makes it a lot more um flexible as far

as use. Deep learning forj is Java based

and then cross is just exploding right

now. And this is interesting. Cross is

uh working with TensorFlow. It actually

can sit on top of TensorFlow. And it can

also do its own thing. Uh so if you're

studying deep learning, you're getting

into it, you want to know the basics of

TensorFlow, but you also are going to

want to know the upper level of KAS

sitting on top of TensorFlow. We're just

looking at TensorFlow today though in

our example. And there's also Torch on

there. There's a bunch more that we

didn't list on here. Um even sklearn or

the uh side package in Python has a

neural network you can program a very

basic one and it is the same basic one

that you could do in TensorFlow if you

stripped everything out of it and then

TensorFlow has a lot of tools they've

added in and so has KAS but we're going

to be looking specifically at TensorFlow

in our example and TensorFlow is an

open- source tool used to define and run

computations on what they call tensors

very common language now so you more and

more we see the term tensor as being a

standard in the uh deep learning

language and this was originally

developed by Google. So let's dig a

little bit big in there. What are

tensors? Tensors are just another name

for arrays. So a tensor of dimension

five. You can see here we have ab kmq

whatever. So it's an array coming in.

And the tensor of dimension 54 more like

a picture. Very common to see that in a

picture. You can also see a tensor even

more detailed than a picture as we go to

the next one. Tensor of dimension 333.

This is 3D space. You might have a

picture that also has colors. That might

be the third dimension. You might have

four dimensions because you have both

your grid and your different color

channels and your zplot. You can see

where you can now process a very

highlevel set of data coming in whether

as an image or features. They could be

features that have nothing to do with

images. So there's a lot of stuff you

can do now with the tensors coming in.

Thus this where the term tensorflow

comes from. So we have um right now the

TensorFlow is the most popular library

in deep learning and I did mention KAS

now works with TensorFlow. So there's a

lot of stuff you can do between the two.

Uh it's an open-source software library

developed by Google. Uh so they hit a

roadblock and they realized hey this is

an infant stage technology. You know we

thought it was going to be the next

greatest thing and we were going to have

a hold on it but it's really infant as

far as how it's applied and what we can

do with it. Let's open source it so

everybody can work on it. uh let's take

it to the next level. And that's really

what open source does to a lot of these

uh packages when they release them. And

you can run on either a CPU or a GPU. So

when we look at the details, if you have

your graphic processing units, um what's

nice about those is they run a lot

faster. The downside is you have to play

with them a little bit to get them up

and running. And it's a hardware

upgrade. When we run it, I'll be running

it in the CPU mode. I have played with

it in my GPU on my personal computer.

you know, it does increase the

processing. Uh, but I did run into some

version problems with my Python and

stuff like that. And when I did finally

work it out, I went back to the CPU

because it didn't increase my speed

enough for what I was working on. But in

a larger group, you might be able put

that on. If you're working with a larger

stack of computers, you might want to

run it in the GPU. You can create a data

flow graphs that have nodes and edges.

So there's our edges coming in. We

didn't talk about edges, but that's very

up and cominging way of looking at your

analytical data is how do different

nodes connect? What do those edges look

like in between them? And it's used for

machine learning applications such as

neural networks. It is mostly a neural

network, but they have all kinds of

tools which sit on top of our basic

neural network. They have new stuff

evolving into the TensorFlow library.

So, it's very much uh just exploding.

great time to jump into TensorFlow

because there's all kinds of cool things

we're doing with it and all kinds of

cool applications you can now use uh

TensorFlow for. So let's take a look at

implementation in TensorFlow and we're

going to build a neural network to

identify handwritten digits using the uh

Mnest database or the MNIST database and

that stands for modified National

Institute of Standards and Technology

database. It is a collection of 70,000

handwritten digits and the digit labels

identify each of the digits from 0ero to

nine. This is a cool example because

it's simple enough that you could

actually run this through some basic

machine learning categorizing algorithms

and train them and you'll get about the

same answer because again it's it's

simple grid. The digits on the grid

don't have a huge amount of variation

like you would say an automated driving

car looking at the environment. So you

can still do this with a lot of your um

different linear models and stuff like

that. You can solve this and you'll get

about the same answer. When I ran a

comparison between TensorFlow and

between some basic uh regression models

or category models uh in machine

learning, they came up pretty even as

far as their output. Uh so this is kind

of where we start to see the complexity

of something coming in this case a

tensor you know or a grid of uh

information where the deep learning

model does as good as the regular models

and when you get past this kind of

complexity and features suddenly the

neural networks come up with better

answers better solutions and a better

build and that's why there's such a move

into neural networks is we live in a

complicated world and it's just really

cool we can do with this. So the

handwritten digits from the um NIST

database, they come in, the data set is

used to train the machine, a new image

of a digit is fed and the digit is

identified. Um and if you've looked at

any of our other machine learning tools

where we're doing training, uh where we

train our uh model to fit and then you

test it out, this should look pretty

familiar. Uh and there is some tools out

there for say untrained categorizing uh

where it's just looking for features

that fit together. So there are tools

that don't need that training. But this

is where uh when we talk about neural

networks, we do need to train them. And

this is what we're looking at.

So for this I'm going to use the

Anaconda Navigator just because it's a

very nice visual tool. You might be in

PyCharm or one of your other IDEs for

editing Python because we are looking at

Python TensorFlow. And under Anaconda,

we have the notebook, which is something

we use pretty regularly. And they have

the Jupyter Lab. The Jupyter Lab is the

Jupyter notebook, but with tabs and a

few new features. So, we'll be using the

Jupyter Lab today. And under the

environment, you'll want to go ahead and

and uh if you haven't yet, uh you'll see

that I have a number of different setups

in here. Right now I have the Python

version 36 and the TensorFlow. In this

case I have TensorFlow 1.12. If we

scroll down you can see that uh here we

go. TensorFlow and it's version 1.12.

And in here if you haven't yet you'll

need to install those and go in and just

open our terminal. And u if you've never

used the Anaconda or if you're in your

other thing you might have something

simple like pip. Is what I use for my

install. And you can simply do install

TensorFlow. And that should bring in the

most current version. Now, when I

installed this a few months ago, Python

version, I'm not going to run this

because I already have installed on

here. Python version 3.7, the newest one

out, still had a couple glitches with

the TensorFlow. I believe they've fixed

it as of writing of this, but um I'm

going to stick with 3.6 just so I don't

get any surprises on there. So, this is

Python version 36 with TensorFlow 1.12

on here. And if you haven't installed it

yet, you also want to install Numpy for

this example. That's Numbers Python or

uh nu py. You can just simply run an

install on there. Keep in mind if you're

in Anaconda uh and you've created one of

these environments specific to this,

keep withd.

If you're going to use pip, keep with

pip. Don't install one package with pip

and one under because that's how they

track those version numbers and how they

fit together and you can end up with a

problem. they don't pip doesn't see cond

and vice versa. Uh so just keep that in

mind when you're running your installs.

We'll go ahead and open up Jupyter Lab

and we're going to launch that. So

here's my Jupyter Lab. One of the really

cool features of Jupyter Lab is you have

tabs now. So you can open up multiple uh

notebooks. And this is nice cuz I have

my notes I'm working on and then our

actual window we're looking in. And

we'll go ahead and zoom in a little bit

here. There we go. So you have a nice u

hopefully easy to see fonts. And then

we'll go ahead and do a simple or get

our imports out of the way. Um, and so

we're going to import our TensorFlow as

TF. Uh, that's pretty much a standard

for TensorFlow, numpy, our numbers

Python as py, and we'll import our matt

plot library as plt. Again, these are

very common. So if you see TF or py or

plt, this is a standard that most people

use. Do you have to? No, you could just

do import numpy instead of doing as py.

And then from

tensorflow.acamples.tutorial

tutorials. This is always nice because

they actually include data set we're

going to play with. So, we're going to

import input data. So, there's our data

coming in. That's all we're doing is

telling it this is where it's coming

from. And if we're going to tell where

it's coming from, we need to go ahead

and create a variable with that

information in it. And we'll just call

this uh mnist

or minced. You know, I don't really know

how they pronounce that. I should

probably look that up. It's a very

common data set to use. And there's our

input data. And we're going to read data

sets. And this is um if you look at

this, we imported input data from our

TensorFlow. And so this is a TensorFlow

read statement for their tutorials. So

this isn't like some special Python

setup. This is just their setup. Makes

it easy to pull it in. So once we get

into their data sets, we need to go

ahead and tell it what kind of data set.

And again, this is what we brought in,

but it's going to be the nint data. And

this part is very important. one hot

equals true. This means that instead of

importing a value from 0 to 9, we

evaluate the data set. It's going to

bring it in as one hot. Whenever you see

one hot encoder, we're flattening that

out. And we have true false for zero,

true false for one, true false for two.

So our output, if you remember from our

output uh from the slide we did earlier,

uh in this case, I grabbed the one for

bike price. Doesn't really matter which

one we use. This has one output. So we

have our bike price on this. We're going

to have instead of one output, we're

going to have 10 outputs representing

each of the digits in there. And this

code really isn't going to show us

anything. It's good to see what we're

actually looking at. So um let's go

ahead and do a figure ax equals go into

our plot library subplots 10, 10. And

that is if you remember we talked about

tensor. Tensor being data coming in.

This is a 10 by 10 grid or 100 pixels on

there. And if we're going to display it,

uh let's go do K0

for I and range 10. Just a simple loop

through on the data. Let's do what is

it? Uh for J and range 10. And I

actually misqued that 10 uh 10 x 10 is

not the actual size of the pixels. Uh

the actual pixels are going to be um we

look at the shapes and we'll get into

that in just a second here. We'll take a

quick look at shape on there. Uh turns

out they're uh what are they? are, I

believe, 28x 28. Uh, so let's take a

look at that. And we're just going to

plot these. What are we looking at? What

are we working with? As a data

scientist, you should always be looking

back at your data and seeing what it

looks like and get that human

perspective because you just never know.

You know, the the computer may put

something out that looks makes no sense.

And at that point, you want to go back

and reevaluate what you did. Uh, so

we're going to go ahead and plot. We're

going to plot 10 digit, you know, 10 of

the digits by 10 of the digits. And

here's our ax. We'll create the J on our

subplots and we're going to do an image

show. We're going to look at the

training image for images of K. And then

we want to reshape this. We're going to

reshape this. And we're going to reshape

this 28x 28. That's how I knew I had it

wrong is cuz I looked down my notes. I

was like, oh no, that says 28. It's not

10 x 10. And I should know that already

cuz I've done enough messing with this

data set that I should have remembered.

Uh, but it's 28 x 28. And the aspect

we're going to do is auto. And this is

all, if you look at this, here's our

variable NIST. the NIST is coming from

data set. Uh so this is all part of the

TF TensorFlow learning or examples

tutorial in there. And then we'll go

ahead and do K plus equals 1. So we just

keep paging through our different um

images. And let's see what that looks

like. Let's go ahead and do a plot show.

Uh and we'll go ahead and run this so we

can take a look and see what we have

here. And so we have a nice plot here.

And you can just see that we have uh

some random numbers showing up in each

one of these little subplots. If you're

wanting a copy of this code, put a note

down in the YouTube video and let us

know or come visit us at

www.simplearn.com

and we'll send you out a copy of what

we're working on and get a copy of that

for your own setup. Uh so now we've

taken a look and we can just see we have

here's our pictures that are coming on.

We plotted them so we have an idea of

what we're looking at. Let's go ahead

and uh print. Let's look at the shape of

the features. Uh so when we have this we

have our nest train images and we'll do

the shape on there. Let's take a look

and just see what we're looking at uh as

far as uh our count and everything. And

so you can see here we have 55,000.

That's basically how many images we have

and this by 784. And in this data set

there's also our labels. So let's take a

look at that. We have our net train

labels shape. Let's take a look and see

what that looks like. Uh and there we

have 10 because there's 10 digits. So we

brought in that's our output we're

looking at. And so we have there we go

55,000. They match. They should match

because you should have equal numbers in

both of those. You know, here's our data

in and here's our answer. If you

remember, this is a bunch of zeros and

with one each each one will be 0001

would be what letter four or something

like that. So, let's take a look at what

our one hot encoding did for the first

observation. And this is when we're

exploring data, you really want to dig

in there and just see what the heck am I

looking at. So, we're going to look at

the labels. And this would be the first

label that comes up. And we'll go ahead

and run this. And we look at that. You

can see this is what I'm talking about.

0 0 or 1 is 0 2 is 0 3 is 0 four is 0 5

is 0 6 is 0. 7 equals 1. So our very

first label is a seven. But our very

first label comes up that it's a seven.

And so we don't have like 0 through 9.

We have a bunch of zeros and just the

one to mark it as a seven on here. So

now we've kind of looked a quick look at

the data. And in here you might ask some

questions like what is 784? 24 * 24.

Remember that's the size of our grid on

there or our tensor coming in. So 784 is

a setup on there. And we've gone through

all this viewing the data. We'll go

ahead and start looking at our

tensorflow. So let's take our X

variable. This is going to be our

training set. We'll do a placeholder and

then we're going to have these come in

as float. Now if I remember correctly,

they're actually, you know, zero or one

for the values because they're either

but we have them coming in as a float

value. And we have a little bit of a

shape coming in here. And there's our

784. Uh so we let it know that this is

what's what our input is for our

TensorFlow. And this is our training

set. So we'll just put a label on there

to help us uh track that train set. And

then W. And with W, we'll go ahead and

do TF variables. And we'll do this as uh

zeros variables TF zeros. And we'll set

this as as 784 by 10. 10 being the

output. 784 being our number of

variables in and this is our weights.

Remember we have a bias in there too.

And I'll go back over this in just a

second as we see how that fits together

in our tensorflow. And we'll do this one

um with our variables again. We have 10.

So we're going to do the bias. We're

going do it the same kind of format and

setup on here. And so we'll do that as

as TF zeros of 10. So we'll just create

an array of 10 there. And this is our

bias. So with these three lines um and

there's actually they're coming out with

the eager execution which would bypass

some of what we're doing. But this is

important to understand is the first

thing you have to do with TensorFlow is

we have to allocate a space for the

variables and our TF placeholder and our

TF variable with our weights and our

biases. This actually hasn't done

anything yet. So all it is is

placeholders. That's why it's okay to

use zeros. Um you could have just as

easily used ones or anything else and it

wouldn't matter. The next stage is to go

ahead and set up some of the functions

going on. But before we do that, just

note that this hasn't done anything.

Even if I execute it, all it's done is

created placeholders until we do the

final initialization. And so we need to

go ahead and set up. We'll do y=

tf.n.oftmax.

And the code for this is tf.mmoxw.

And this is our uh sum. Let's just put a

note here so we can keep track of what's

going on. We're finding weighted sum of

inputs plus the bias. Uh so there's our

plus b the bias and then we need to go

ahead keep um let's do y underscore and

again another placeholder and this one

we'll set um it actually we'll put in as

tf placeholder on here tf placeholder

float none 10. There's our one hot

encoder going on there. So our 10 values

coming out and we'll do a cross entropy

on here and this is going to be minus tf

reduce sum and we'll do y here's our y

underscore which is remember we have

your y output and your actual output. Uh

so this will be our y underscore time

the tf log of y. And then finally um

before we do the actual initialization

of all our variables we'll set up our

train step. This equals our gradient

descent optimizer. Very important.

Remember we looked at that chart on our

um uh slides and so we've set up all

these formulas and here's our gradient

descent optimizer and as it keeps

looking it keeps looking for that zero

value. That's what we're doing with that

particular formula. So let's take a look

and see what we're doing here. We just

put together all of our pieces for

TensorFlow. And you know the devil's in

the details. We have here our training

set coming in. We have to put a

placeholder on there. We have our uh

variables with their weights. We have

our biases coming out and then we put in

our uh the weighted sum. So here's

summizing our summation here. Then we

have our y variable output. So there's

our y um how it works and then of course

the actual output on there. And then we

have our cross entropy coming in and

that's our minus tf.reduce sum the y *

the tf log of y. And then the training

step gradient descent optimizer and

we're using a 0.01 in this and we're

going to minimize cross entropy. So,

we're going to let it do all the work.

So, once we've set up all of these

different layers, we've allocated for

them, we need to go ahead and initialize

them. So, we're going to do an init tf

initialize, and it's going to be all

variables. Uh, one of the cool things is

they're in the process of doing away

with this. So, all these steps would be

bundled into one instead of having to

have placeholders. You initialize them

in the same process going on. And then

finally, everything in TensorFlow is

based on your session. Now, this is

changing that there's other options to

be able to run this, but we want to go

ahead and do uh session. There's our TF.

There's a TF session. And then we want

to go ahead and do session run. And what

are we going to run? Well, we did

initialization of all our variables. Uh

so, this is what we're running. And this

is we're actually once we do this, we

actually create our TensorFlow object.

So, this whole piece of code right here

is our TensorFlow object. We have our

input coming in with our weighted

variables coming in. Our soft max for

our metal going out. How does it add it

together for our y value? Uh and then we

have the actual uh float value coming

out. Checking on our all the way down.

So you can see all the different stages

going through that we're setting up. Um

and this is one of the reasons that a

lot of people like TensorFlow is because

you can designate all these different

pieces one step at a time. This is also

one of the reasons people don't like

TensorFlow is because you have to

designate all the different layers

coming down and there's a lot of steps

being made right now to minimize this to

make it either easier to automate it or

to allow you to do more complicated

things and all those steps are still at

play. So it's worth looking into the

more advanced version what's going on

with KAS on top of TensorFlow. It's also

important to understand what's going on

in these individual levels if you're

going to play with them. It's important

to understand, hey, what's going on with

the uh finding the weighted sum of the

inputs plus the bias because there's

other ways to do that. There's all kinds

of other tools in there now, but this is

the basic setup that you want to do on a

TensorFlow coming in. And we want to go

ahead and just run and admit our

TensorFlow. So, let's go ahead and do

that. Let's run this. We do get a

warning here because uh there's a move

to use global variables. This is one of

the changes they're making, but it as

far as this example, it's not going to

make a difference because we're doing

once we initialize it. This is

initializing our variables. And again,

these are only placeholders up here

until we initialize them. And I would

highly suggest put a note down there or

or go over to simplylearn.com and let

them know and have them email you a copy

of the code. So, you can actually play

with this code right here because this

is the body of what's going on in

TensorFlow. This is the build in neural

networks. And then once we've done that,

now comes kind of the fun part is we

need to go ahead and train it. Uh so

we've created our TensorFlow, we've

created our uh network and now we need

to go ahead and train it. Uh so let's

put together that training code and

let's just do uh for I in range u 0 to

1,000. So we're just going to look at uh

the first 10,000 in our training. And

the way we pull that data from our mints

train next batch of 100. Uh so you look

at this. We're going to be doing groups

of 100 and then there's going to go

through a thousand of them. This is very

important that TensorFlow builds this

in. This is one of the downsides of

doing sklearn or one of the older

packages is they don't let you batch

groups in. Uh they wanted to have it all

up front and then you have to build your

own batch programs right now. Uh this

lets us go ahead and do that. And you

can see here we have batch x of s, batch

y of s. So there's our x and our y. You

could look at this as our training of X

and our train of Y or the data N and the

answer in. Uh and then we simply do our

session run. Uh so here's our session

that we've created. We're going to run

it and we want to do the train step. We

initialized our train step up here and

our TF. And so there's our train step

feed. It's a dictionary. Dictionary

coming in which we're going to create

right here. uh is x is our batch x of

our sample comma and our y underscore is

going to be our batch of our y sample.

Uh and so this goes through and we've

now hit the run button and we've trained

our session. We've trained this setup on

here. And once we've trained it, then we

need to go ahead and find out how good

our accuracy was and actually start

running some predictions through there.

Uh so we'll go ahead and create a a

correct prediction. And this is where

our tf.equal equal. We'll use our argmax

y of one and tf argmax of y of

underscore of one to help us get the

correct predictions on there. And then

we want to use that to feed into an

accuracy. And so our accuracy is going

to be tf reduce uh mean and we'll take

that and we'll do um a cast and this is

the correct prediction that we're

sending in there. And it is a uh float

value. Keep it simple. And let's go

ahead and print this out so we can see

what we're looking at. Uh so what are we

printing out? Uh we need to do a session

run. This session run is going to be on

the accuracy. Where did accuracy comes

from? This is our we're casting our TF

on there with the correct predictions on

that. So here's our accuracy feed in. So

it needs a dictionary for the data

coming in. We're going to create our

dictionary and x is going to be our nest

test images and y there is going to be

our nest.est

labels. Let me just double check and

make sure I have that typed in there

correctly. There we go. Oh, and let's go

ahead and run that and see what comes

up. And we end up with a N165

for our accuracy, which means our

trained neural network does a pretty

good job letting us know what these

different symbols are in guessing that a

seven and a three and a four, uh,

something that as humans we kind of take

for granted. I even have trouble reading

this. So, I don't know if I would be

able like that first one, I would sit

there for a long time figuring out

that's a seven versus a two. That could

have easily been a two to me. Sing seem

to do a pretty good job analyzing this

data. And this is used to analyze

something very complicated on these

images, very different than uh just a

straight value of uh cost of sales and

here's our return and our marketing. Uh

we can now create this nice neural

network that does all kinds of cool

things. Do you know friends that

according to the lending statistics the

demand for AI and ML specialist is

projected to surge by 40% between 2023

to 2027.

And on an average, an ML engineer is

expected to earn around 133 and $336 per

year. So if you are an aspiring ML

engineer and thinking about what

innovative projects you can show in your

portfolio, then your wait is over cuz in

this video I'll be covering eight

amazing ML projects that you can

showcase in your resume. So guys, let's

start first with a beginner level

project and the first project that we

are going to encounter that is home

value prediction. So guys, this project

aims to develop a predictive model to

estimate the value of residential

properties. The model will analyze

various features such as location,

square, footage, number of bedrooms and

bathrooms, age of the property and other

relevant factors. By leveraging

historical property data, the model will

be able to provide accurate home value

predictions which can be useful for real

estate agents, buyers and sellers. So

guys, the programming language that we

are going to use all over here will be

Python and machine learning libraries

that we will be using will be

scikitlearn, tensorflow, kas and for

data handling libraries we have pandas,

numpy and for visualization we have to

use mattplot and seabon. Now what will

be the approach for this one guys? So

guys the first one that we have a data

collection. So here what is going to

happen guys? So first you have to

collect the historical property data

from the sources like Zillow

retailer.com. You can also get database

from the public real estate databases

like Kaggle data sets where you have

Zillow home value prediction. Ensure

that the data set include features like

location where you have latitude,

longitude, square footage, number of

rooms, year built, property type and

previous sales. The next step that comes

is data cleaning. You have to handle the

missing values by using imputation

techniques or removing incomplete

records. Removing outliers that may skew

the model's prediction, normalize or

standardize the data to ensure

consistency. The third one that we have

is feature engineering. You have to

create new features such as proximity to

schools, crime rates and access to the

public transportation. Encode categorial

variables, example property type,

location using techniques like one hot

encoding. Generate interaction features

that capture relationship between

existing features. The fourth one that

we have is model selection. Use

regression models like linear

regression, random forest, gradient

boosting, neural networks. Experiment

with different models to identify the

best performing one. Now in the next

phase all you have to do guys is model

training and evaluation. Split the data

set into training and test sets. Train

the model on a training set and evaluate

their performance on the testing set

using metrics like RSM which means root

mean squared error. You can use cross

validation to ensure the model's

robustness and avoid overfitting. The

sixth one that we have all over here is

hyperparameter tuning. You can optimize

the model's hyperparameter using

techniques such as grid search or random

search to improve accuracy. And if

you're looking forward to deploy your

model, then you can develop a web

interface using flask or Django to allow

users to input property features and get

predictions. You can deploy the model on

the cloud platform like AWS for

scalability.

Now if we talk about the complexity

level of this, we all know that it is a

beginner level project. Now let us move

on to the one more set that is music

genre classification and generation. So

guys this is also one of the most

beginner level project. This project

aims to develop a system that can

classify music tracks into different

genres and generate new music

composition within specified genre. The

goal is to build a model that analyzes

audio features to categorize music and

uses deep learning techniques to create

new music. This project introduces

advanced concept of audio processing,

deep learning and generative models. So

guys, what will be used in this? So

we'll have programming language that

will be Python. For audio processing,

we'll be using librosa. For machine

learning libraries, we'll be using

tensorflow, kas, pytorch. For data

handling libraries, we'll be using

pandas, numpy. For visualization, we'll

be using mattplot, seabon. For the data

set guys, you can use gtzan music genre

data set or you can get it from free

music archive. So guys in the first

phase we are going to have data

collection. You can obtain data sets

containing music tracks and their

corresponding genre from the sources

from GTN music genre data set and the

free music archive. Ensure that a data

set includes diverse genre and

substantial number of tracks per genre.

Next we'll go for data prep-processing.

Use library librosa to load and

pre-process audio files including

feature extraction such as mil

frequency, septal coefficients, chroma

features and spectral contrast. You can

normalize the extracted features to

ensure consistent input for the given

models. Now if you talk about feature

engineering guys, you can extract

additional features from the audio files

such as tempo, beat, zero crossing rate

etc. Create a feature matrix that

represents the extracted audio features.

Then go for the model selection. Use

conventional neural network or recurrent

neural networks for the music genre

classification. Split the data set into

training and testing data set. Now if we

talk about model training and evaluation

guys then you can train the selected

classification model on the training

set. Evaluate the model's performance on

the testing set using metrics like

accuracy, precision, recall and F1

score. Use confusion matrices to

understand the classification

performance across different genres. Now

if we talk about model selection and

training for the music generation, what

you will do guys? You can use the

generative adversial networks or

recurrent neural networks such as LSTM,

long short-term memory for music

generation. Train the generative model

on the data set to create new music

sequences. Next, we have model training

and evaluation. You can train the

generative model on sequences of audio

features. You can evaluate the generated

music by listening tests and by

objective metrics like inception score

or fche audio distance. If I talk about

hyperparameter tuning guides, we can

optimize the model. You can use the

hyperparameters using techniques like

grid search or random search to improve

performance. If I talk about deployment

guys, you can deploy these models on

cloud platforms like AWS. Now let us

move on to our next project. So guys,

the complexity of this project is at the

beginner level. Now let us move to the

intermediate level projects. Next

project that we have all over here is

sentiment analysis of Twitter data. This

project aims to develop a sentiment

analysis model that can classify to its

side positive, negative or neutral. The

goal is to analyze public sentiment on

various topics or events using natural

language techniques. So guys, what will

be used all over here? So in this we

will have programming language like

Python. Okay. NLP libraries, NLTK spacy.

For machine learning libraries you can

use scikitlearn, tensorflow, kas. For

data handling libraries we have pandas,

numpy. For visualization we have

mattplot lilip seon and we can use the

API Twitter for data collection. Now how

you going to work on it guys? So guys if

I talk about the data collection use a

Twitter API to collect tweets based on

specific hashtags like keywords or

topics. Extract relevant fields like

tweet text, user information, timestamp

etc. Then if I talk about data

prep-processing guys, clean the tweet

text by removing special characters,

links, mentions, hashtags and stop

words. Tokenize the text and perform

limization or stemming to reduce the

words to their base form. Next, if you

talk about feature engineering guys, you

convert the clean text data into

numerical representation using TF, back

of words or word embedding. Now if I

talk about model selection guys, you can

choose a classification algorithm such

as logistic regression, n bias or LSDM.

Split the data set into training and

testing data set. Now if I talk about

model training and evaluation, then you

can train the selected model on the

training set. Evaluate the model's

performance on the testing set using

metrics like accuracy, precision,

recall, and fn score. Use cross

validation to ensure the model's

robustness. If I talk about

hyperparameter tuning guys, you can

optimize the model's hyperparameters

using grid search or random search to

improve the performance for deployment

which can be optional. You can deploy

your model on AWS for real-time

sentiment analysis. So if I talk about

the complexity level guys, its

complexity is intermediate. So guys, our

next project is customer segmentation

using K means clustering. This project

aims to segment customers into distinct

groups based on their purchasing

behavior and demographic information.

The objective is to understand customer

segments and tailor marketing strategies

accordingly. So guys, what programming

languages we'll be using? So basically

we'll be using Python. For machine

learning libraries, we will have

scikitlearn. For data handling

libraries, we'll have pandas, numpy. For

visualization libraries, we'll have

mattplot lab, seabon. And the data set

source will be e-commerce transaction

data. how we are going to work on this

one. For data collection, we can obtain

a data set of e-commerce transactions

that include customer demographics,

purchase history, and product

information. Next, we'll have data

prep-processing. For data

prep-processing, we are going to do the

cleaning of the data by handling the

missing values and outliers. Then, for

feature engineering, we are going to

create features like total purchase

amount, purchase frequency, and recency

of the purchases.

Then, we are going to proceed for the

model selection. You can use K means

clustering to segment the customers into

distinct groups. You can determine the

optimal number of clusters using methods

like album methods or silhou.

Now if I talk about model training and

evaluation, you can train the K means

model on the process data set. You can

evaluate the quality of clusters by

analyzing intracluster and intercluster

distances. Next we have the evaluation.

You can visualize the clusters using

techniques like PCA, principal component

analysis, TSN etc. Next we have the

hyperparameter tuning. Now now you can

tune this model and interpret the

characteristic of each segment. You can

develop a target marketing strategies

for each segment based on unique

behavior and preferences. Now deployment

is optional. You can develop a dashboard

using flask or Django to visualize

customer segments and track marketing

campaigns. So guys if I talk about the

complexity of this project. So this is

an intermediate level project. So guys

for data set you can use the Kaggle's

customer segmentation data set which is

available at the Kaggle's platform. Now

the third intermediate level project

that we have all over here is building a

chatbot with Rasa. This project aims to

build an intelligent chatbot using Rasa

framework. The chatbot will be capable

of understanding user queries and

providing appropriate responses making

it useful for customer support, personal

assistance or information retrieval.

What languages we are going to use? So

it will be Python based. We'll have the

NLP libraries like Rasa, NLTK, Spacey.

For machine learning libraries, we are

going to have scikitlearn, tensorflow,

kas, etc. For data handling libraries,

we are going to use pandas, numpy. So

guys, this was what we are going to do

it and how you can work on this one by

collecting the data, collect the

conversation data and FAQs from the

target domain, annotate the data to

create training examples for the

chatbot. Next comes is data

prep-processing. Clean the text data by

removing special characters and

normalizing the text. You can tokenize

and limitize the text to prepare for a

training. The third one we have the

models training. You can use Rasa's

NLU's component to train a model for

intent recognition and entity

extraction. You can define a dialog

management policies to handle different

conversation flows. Next guys, you can

perform the feature engineering and

integration. You can integrate the Ras

NLU and core components to build

complete chatbot. You can connect the

chatbot to messaging platform like

Facebook Messenger etc. For model

selection and testing what you can do

guys you can test this chatbot with

various inputs to ensure that it handles

the scenarios appropriately and you can

also select the right model using this.

Now if I talk about model training guys

what you have to do you have to collect

the user feedback and conversational

logs to continuously work on training

the model. Next, similarly you have to

retrain the model periodically with the

new data to see how it is working. So

that will be your evaluation. Now for

the hyperparameter tuning, what are you

going to do guys? You have to check in

those scenarios where it is able to tune

up with those scenarios where it can

handle the input appropriately. And next

is deployment. So guys, for deploying

it, you can use AWS. So guys, for a data

set, you can use Ras open source. So

that's a very good data set for you to

proceed. So guys, if you talk about

difficulty of this project, this is an

intermediate level project. Now let us

move on to the advanced level projects.

For advanced level projects, the first

one that comes up to my mind is movie

similarity from plot summaries. Now this

project aims to develop a system that

can find out recommended movies similar

to a given movie based on their plot

summaries. By analyzing the textual

content of the movie plot summaries, the

model will identify similarities and

suggests movies with similar themes,

story lines or genres. This project

introduces beginners to natural language

processing text similarly measures and

recommendation systems. What languages

we are going to use guys? We'll be using

Python NLP libraries like NLTK spacy

machine learning libraries like

scikitlearn data handling libraries like

pandas numpy visualization we can use

numpy and data set source will be IMDb

or kegel so guys this process is also

involving the data collection then you

have to go for data cleaning then

feature engineering next model selection

so similar process as I have discussed

in other projects so you have to also go

through the same one next what you have

to do device. Similarly, what you have

to do, you have to train the model, then

evaluate the model, then hypertune it

and finally proceed for the deployment.

So, this is overall process of this

project. Try to research on the website

a lot like how you can extract it. So,

guys, you can use kegel or towards data

science to research more about this

project. Now guys, if I talk about the

difficulty of this project and this is

an advanced level project. Now let us

move on to the next one that we have all

over here that is image segmentation

project for brain tumor prognosis. This

is a very very amazing project and

definitely you can put up on your

portfolio. Basically guys this project

aims to develop an image segmentation

model to identify and delinate brain

tumors from MRI scans. The goal is to

accurately segment the tumor regions

which can aid in prognosis treatment

planning surgical interventions. This

project introduces intermediate level

concepts of computer visions, deep

learning and medical image analysis. So

guys, what we'll be using all over here

for programming languages we can use

python for deep learning libraries we

can use tensorflow, kas, pytor. For

image processing libraries, we can use

opencv, scikit image. For data handling

libraries, we can use pandas, numpy. For

visualization, we can use mattplot,

cbond. Now if I talk about what is the

process of developing this project the

first step will be same data collection.

So next step you have to go for data

prep-processing. Third step you have to

do the model selection where you can use

CNN models for image segmentation task.

Then you go for model selection. Moving

ahead you're going to have the model

training and evaluation. You have to

split the data set into training and

validation and testing data sets. Next

proceed for the evaluation phase. Okay,

evaluate the model with certain metrics.

So here I can give you certain idea like

you can use dice coefficient,

intersection over union or accuracy.

Then go for hyperparameter tuning where

you have to optimize the model's

hyperparameters.

You can use grid search or random search

as we have discussed. And finally you

can deploy this model on AWS. Now guys

we have come to the final project. This

is also a very amazing project guys. So

guys the complexity level of this

project is advanced level.

Now let us move on to our final project

that is the impact of climate change on

birds. This is a very very amazing

project and definitely you can add it on

your resume. This project aims to

analyze the impact of climate change on

the bird population and migration

patterns by examining various climatic

factors and their correlation with bird

species data. The project seeks to

predict how climate change might affect

bird behavior and distribution. This

project will introduce you some advanced

level concepts like time series

analysis, environmental data modeling,

etc. So guys, what programming languages

we'll be using for data analysis? You

can see we'll have pandas, numpy. For

machine learning libraries, we are going

to have scikitlearn, tensorflow. For

visualization, we're going to have

mattplot lab, plotly. For geospatial, we

are going to have geopandas, folium. For

data source, we're going to have public

data sets on bird observation and

climate data sources from eird. Now what

will the process flow for this one guys?

First you have to proceed for data

collection. Gather bird observation from

the data like EIRD which provides

extensive record of bird sightings.

Okay. And for climate data you can

collect it from NOA including

temperature, precipitation and other

relevant climatic factors over the time.

And similar next process will be the

data prep-processing. Then you have to

proceed for feature engineering. Then

you have to go for model selection.

Okay. Moving ahead you have to go for

model training. then evaluation, then

hyperparameter tuning and finally you

have to deploy the model. So research

about this project, see what models you

are going to use. Suppose I can give you

a hint about this. You can use time

series analysis models like ARMA or ML

models. You can also use random forest

or gradient boosting for predicting

impact on the bird population. So guys,

use Google exhaustively to research

about this project. This is also a very

amazing project and it's going to give

you a lot of idea. Now if I talk about

the complexity level of this, it is an

advanced level project.

>> Welcome to deep learning interview

questions. My name is Richard Kersner

with the SimplyLearn team. That's

www.simplearn.com.

Get certified, get ahead. Today we're

going to help you prepare for interview

questions dealing with deep learning.

And we're going to go from the very

basics of neural networks and deep

learning into some of the more commonly

used models so you can have an

understanding of what kind of questions

are going to come up and what you need

to know in interview questions. We'll

start with a very general concept of

what is deep learning. This is where we

take large volumes of data in this case

on cats and dogs or whatever. A lot of

times you use um a training setup to

train your model. Remember it's kind of

like a magic black box going on there.

And then we use that to extract features

or extract information and in this case

classify the image of a cat and a dog.

So the primary takeaway we're talking

about deep learning is it learns from

large volumes of structured and even

unstructured data and uses complex

algorithms to train neural network. It

also performs complex operations to

extract hidden patterns and features.

And if we're going to discuss deep

learning in this very uh simplified

overview and we also have to go over

what is a neural network. This is a

common image you'll see of a drawing of

a forward propagation neural network and

it's it's a human brain inspired system

which replicate the way humans learn. So

this has inspired how our own neurons

and our brain fire but at a much

simplified level. Obviously it's not

ready to take over the human uh

population and and be our leader yet.

Not for many years. It's very much in

its infant stage. But it's inspired by

how our brains work. Um and they use a

lot of other inspirations. You can study

brains of moths and other animals that

they've used to figure out how to

improve these neural networks. The most

common one consists of three layers of

network and this is generally how you

view these networks is you have an

input, you have a hidden layer and an

output. And the neural network is uh

broken up into many pieces. But when we

focus just on the neural network, it's

always on the hidden layers that we're

making all the adjustments and figuring

out how to best set up those hidden

layers for their functions to both train

faster and to function better. When we

look at this, of course, we have our

input, hidden, and output. Each layer

contains neurons called as nodes perform

various operations. And you can see here

we have the list of the nodes. We have

both our input nodes and our output

nodes and then our hidden layer nodes.

And it's used in deep learning algorithm

like CNN, RNN, GN, etc. We'll address

some of these models a little closer, at

least the most common models as we go

down the list and we study the deep

learning and the neural network

framework. Let's start with what is a

multi-layer perceptron or MLP a lot of

time as they're referred to. And you'll

see these abbreviations. I'll be honest,

I have to write them down on a piece of

paper and go through them because I

never remember what they all mean even

though I play with them all the time.

What is a multi-layer perceptron? Well,

if you look at the image on the right,

it's very similar to what we just looked

at. You have your input layer, your

hidden layer, and your output layer. And

that's exactly what this is. It has the

same structure of a single layer

perceptron with one or more hidden

layers except the input layer, each node

in the other layers uses a nonlinear

activation function. What that means is

your input layer is your data coming in

and then your activation function is

based upon all those nodes and weights

being added together and then it has the

output. MLP uses supervised learning

method called back propagation for

training the model. Very key word there

is back propagation. Single layer

perceptron can classify only linear

separable classes with binary output 01.

But the MLP can classify nonlinear

classes. So let's break this down just a

little bit. The multi-layer perceptron

with an input layer and a hidden layer

and an output layer. As you see that it

comes in there, it has adds up all the

numbers and weights depending on how

your setup is. That then goes to the

next layer. That then goes to the next

hidden layer if you have multiple hidden

layers. And finally to the output layer.

The back propagation takes the error

that it sees. So whatever the output is,

it says, hey, this has an error to it.

It's wrong. And then sends that error

backwards from where it came from. And

there's a lot of different functions

used to uh train this based on that

error and how that error goes backwards

in the notes. Uh so forward is you get

your answers. Backward is for training.

You see this every day. Even my uh

Google Pixel phone has this. It they

train the neural network which takes a

lot more data to train than it does to

use. And then they load up that neural

network into in this case I have a Pixel

2 which actually has a built-in neural

network for processing pictures. And so

it's just the forward propagation I use

when it processes my photos, but when

they were training it, you use the back

propagation to train it with the errors

they had. We'll be coming back to

different models that are used. For

right now though, multi-layer

perceptron, MLP, put that down as your

vocabulary word and of course back

propagation. What is data normalization

and why do we need it? This is so

important. We spend so much time in

normalizing our data and getting our

data clean and setting it up. Uh so we

talk about data there's a pre-processing

step to standardize the data. So

whatever we have coming in we don't want

it to be a uh you know one gigabyte file

here a 2 GBTE picture here and a 3

kilobyte text there. Even as a human I

can't process those all in the same

group. I have to reformat them in some

way that loops them together so they're

a standardized format. We use this uh

data normalization and and

pre-processing to reduce and eliminate

data redundancy. A lot of times the data

comes in and you end up with two of the

same images or uh uh the same

information in different formats. Then

we want to rescale values to fit into a

particular range for achieving better

convergence. What this means is with

most neural networks they form a bias.

We've seen this in recently in attacks

on neural networks where they light up

one pixel or one piece of the view and

it skews the whole answer. So suddenly u

because one pixel is really bright uh it

doesn't know what to do. Well when we

start rescaling it we put all the values

between say minus one and one and we

change them and refit them to those

values. It helps get rid of that bias

helps fix for some of those problems.

And then finally we restructure the data

and improve the integrity. We want to

make sure that we're not missing values

um or we don't have partial data coming

in. One way to look at this is uh bad

data in bad data out. And so you want

clean data in and you want good answers

coming out. One of the most basic models

used is a Boltzman machine. So let's

address what is a Boltzman machine. And

if you know we just did the MLP

multi-layer perceptron. So now we're

going to come into almost a simplified

version of that. And in this we have our

visible input layer and we have our

hidden layer. The Boltzman machines are

almost always shallow. They're usually

just two-layer neural nets that make

stochastic decisions whether a neuron

should be on or off. True or false? Yes.

No. First layer is a visible layer and

second layer is the hidden layer. Nodes

are connected to each other across

layers, but no two nodes of the same

layer are connected. Hence, it is also

known as restricted Boltzman machine.

Now that we've covered a basic MLP or

multi-layer perceptron, and we've gone

over the Boltzman machine, also known as

the restricted Boltzman machine, let's

talk a little bit about activation

formulas. And this is a huge topic that

can get really complicated but it also

is automated. So it's very simple. So

you have both a complicated and a simple

at the same time. So what is the role of

activation functions in a neural

network? Activation function decides

whether a neuron should be fired or not.

That's the most basic one and that

actually changes a little bit because

it's either whether fired or not in this

case activation function or what value

should come out when it's fired. But in

these models, we're looking at just the

boltsman restricted layers. So this is

what causes them to fire. Either they

don't or they do. It's a yes or no,

true, false, all or nothing. It accepts

the weighted sum of the inputs, the bias

as input to any activation function. So

whatever activation function is, it

needs to have the sum of the weights

times the input. So each input, if you

remember on that model, and let's just

go back to that model real quick. And

then you always have to add a bias. And

you can look at the bias if you remember

from your uklitian geometry. You draw a

straight line. Formula for that line has

a y-coordinate at the end. It's always

um cx plus m or something like that

where m is where it crosses the

ycoordinates. If you're doing a straight

line with these weights, it's very

similar, but a lot of times we just add

it in as its own weight. We take it as a

node of a one value coming in and then

we compute its new weight. And that's

how we compute that bias just like we

compute all the other weights coming in.

The node which gets fired depends on the

y value. And then we have a step

function. And the step function this is

where remember I said it's going to get

complicated and simple all at the same

time. We have a lot of different step

functions. We have the sigmoid function.

We have just a standard step function.

We have the ru is pronounced like ray

the ray of from the sun and lu like a

name. So ru function. And we have the

tangent h function. And if you look at

these, they all have something similar.

They all either force it to be um one

value or the other. They force it to be

in the case of the first three a zero or

one. And in the last one, it's either a

minus one or one. And you can easily

convert that to a 0, one, yes, no, true,

false. And on this, one of the most

common ones is the step function itself

because there is no middle value. There

is no um uh discrepancy that says, well,

I'm not quite sure. But as you get into

different models, probably the most

commonly used used to be the sigmoid was

most commonly used, but I see the relu

used more often. Really, depending on

what you're doing, you just have to play

with these and find out which one works

best depending on the data in your

output. The reason to have a non01

answer or something kind of in the

middle is when you're looking at this

and it's coming out, you can actually

process that middle ground as part of

the answer into another neural network.

So it might be that the relu function

says hey this is only a 6 not a one and

uh even though the one is what's going

into the next neural network or the next

hidden layer as an input the 6 value

might also be going in there to let you

know hey this is not a straight up one

or straight up zero it's someplace in

the middle this is a little uncertain

what's coming out here so it's a very

powerful tool in the basic neural

network you usually just use the step

function it's yes or no let's take a um

a big step back and take a kind of an

overview. The next function is what is a

cost function that we're going to cover.

This is so important because this is

your end result that you're going to do

over and over again and use to decide

whether the model is working or not,

whether you need to try a different step

function, whether you need to try a

different activation, whether you need

to try a fully different model used. Uh

so what is the cost function? Cost

function is a measure to evaluate how

good your model's performance is. It is

also referred as loss or error used to

compute the error of the output layer

during back propagation. There's our

back propagation where we're training

our model. That's one of our key words.

Mean squared error is an example of a

popular cost function. And so here we

have the cost function C = half of Y - Y

predicted. Um and then you square that.

So the first thing is um you know real

quick if you haven't done statistics

this is not a percentage. It's not a

percentage of how accurate it is. is

just a measurement of the error and we

take that error if we're training it and

we push that error backwards through the

neural network and we use that through

the different training functions

depending on what model you're using to

train the neural network. So when you

deploy the network you're usually done

training it because it takes a lot of

computational force to train it. Um this

is a very simple model and so you deploy

the train one. Uh but we want to know

how your error is and so how do we do

that? Well you split your data. part of

your data is for trading and part of

your data is for testing. And then we

can also test the error on there. So

it's very important. And then we're

going to go one more step on this. We

got to look at both the local and the

global setup. It might work great to

test your data on what you have on your

computer, but that's different than in

the field. So, when we're talking about

all these different tests and the error

test as far as your loss, you don't you

want to make sure that you're in a

closed environment when you do initial

testing, but you also want to open that

up and make sure you follow up with the

testing on the larger scale of data

because it will change. It might not fit

the larger scale. There might be

something in there in the way you

brought the data in specifically or the

data group you used or um any of those

could cause an error. So, it's very

important to remember that we're looking

at both the local and the global context

of our error. And just one other side

note on a lot of the newer models of

neural networks by comparing the error

we get on the data our training data

with a portion of the test data we can

actually figure out how good the model

is whether it's overfitted or not. We'll

go into that a little bit more as we go

into some of the different models. So we

have our output. We're able to um figure

out the error on it based on the square

means usually although there's other uh

functions used. So we want to talk about

what is gradient descent? Another

vocabulary word gradient descent is an

optimation algorithm to minimize the

cost function or to minimize the error.

Aim is to find the local or global

minima of a function. Determine the

direction the model should take to

reduce the error. So as we're looking at

this, we have our uh squared error that

we just figured out the co based on the

cost function. It says how bad is my

model fitting the data I just put

through it. And then we want to reduce

that error. So how do you figure out

what direction to do that in? Well, it

could be that you're looking at just

that line of that line of data coming

in. So that would be a local minima. We

want to know the error of that

particular setup coming in. And then you

have your global your global minima. We

want to minimize it based on the overall

data we're putting through it. And with

this we can figure out the global

minimum cost. We want to take all those

local minimum costs of each piece of

data coming in and figure out the global

one. How are we going to adjust this

model to fit all the data? We don't want

it to be biased just on three or four

lines of data coming in. We want it to

kind of extrapolate a general answer for

all the data coming in. But this of

course uh we mentioned it briefly about

back propagation. This is where really

comes in handy is training our model.

Neural network technique to minimize the

cost function helps to improve the

performance of the network. Back

propagates the error and updates the

weights to reduce the error. So as you

can see here is a very nice depiction of

a back propagation. We have our

predicted y coming out and then we have

since it's a training set we already

know the answer and the answer comes

back and based on case of the square

means was one of the functions we looked

at uh one of the activation functions

based on cost function that cost

function then depending on what you

choose for your back propagation method

and there's a number of them will change

the weights it will change the weight

going to each of one of those nodes in

the hidden layer and then based upon the

error that's still being carried back

it'll change the weights going to the

next hidden layer and then it computes

an error level on that and sends that

back up. And you're going to say, well,

if it computes the error into the first

hidden layer and fixes it, why would it

stop there? Well, remember, we don't

want to create a biased neural network.

So, we only make small adjustments on

these weights. We don't make a big

adjustment that changes everything right

off the bat. So, no matter how far back

you go, you're always going to have a

small amount of error, and that's still

going to continue to go all the way back

up the hidden layers. For right now,

focus on the back propagation is taking

that error and moving it backwards on

the neural network to change the weights

and help program it so that it'll have

the correct answers. So far, we've been

talking about forward propagation neural

networks. Everything goes forwards, goes

left to right. Uh but let's let's take a

little detour and let's see what is the

difference between a feed forward neural

network and a recurrent neural network.

Now, this is in the function, not when

we're training it using the back

propagation. So, you've got new

information coming in and you want to

get the answer and there's a couple

different networks out there and we want

to know we have a feed forward neural

network and we have a new uh vocabulary

term recurrent neural network. A feed

forward neural network signals travel in

one direction from input to output. No

feedback loops considers only the

current input cannot memorize previous

inputs. One example of one of these feed

forward neural networks. And we've

covered a number of them, but one of the

ones that has a big highlight nowadays

is the CNN, a convolutional neural

network. TensorFlow, the one put out by

Google is probably most known for their

CNN, where the information goes forward.

It uh first takes a picture, splits it

apart, goes through the individual

pixels on the picture, so it picks up a

different reading, then calculates based

on that, goes into a regular feed

forward neural network, and then gives

you a categorization on there. Now,

we're not covering the CNN today, but we

do have a video out that you can look up

on YouTube put out by SimplyLearn, the

convolutional neural network. wonderful

tutorial. Check that out and learn a lot

more about the convolutional neural

network. But you do need to know that

the CNN is a forward propagation neural

network only. So it's only moving in one

direction. So we want to look at a

recurrent neural network. Signals travel

in both directions making it a looped

network. Considers the current input

along with the previous received inputs

for generating the output of a layer.

Has the ability to memorize past data

due to its internal memory. And you can

see they have a nice uh image here. We

have our um input and for some reason

they always do the recurrent neural

network um in reverse from bottom up in

the images. It's kind of a standard

although I'm not sure why. Your X goes

into your hidden layer and your hidden

layer the answer for part of the answer

from that it generates feeds back into

the hidden layer. So now you have an

input of both X and part of the hidden

layer and then that feeds into your

output. Now if we go back to the forward

let me just go back a slide and we're

looking at uh our forward propagation

network. One of the tricks you can do to

use just a forward propagation network

is if you're in a what they call a time

sequence, that's a good uh term to

remember or a time series meaning that

it's sequential data. Each term comes

after the other. You can trick this by

creating your input nodes as with the

history. So if you know that uh you have

values one, five and seven going in and

you know what the output is from one

what those outputs are, you can expand

the input to include the history input.

That's one of the ways to trick a

forward propagation network into looking

at that. But when you do with a

recurrent neural network, you let the

hidden layer do that for you. It sends

that data and reprocesses it back into

itself. What are some of the

applications of recurrent neural

network? The RNN can be used for

sentiment analysis and text mining.

Getting up early in the morning is good

for health and it's a positive

sentiment. One of the catches you really

want to look at this when you're looking

at the language is that I could switch

this around and totally negate the

meaning of what I'm doing. So, it no

longer be positive. So, when you're

looking at a sentence, knowing the order

of the words is as important as the

meaning of the words. You can't just

count how many good words there are

versus bad words to get positive

sentiment. You know, have to know what

they're addressing. And there's lots of

other different uses. Uh, kids are

playing football or soccer as we call it

in the US. RN can help you caption an

image. So based on previous information

coming in, it refeeds that back in and

you have a image setter. And then time

series problems like predicting the

prices of stocks in a month or quarter

or sell of product can be solved using

an RNN. And this is a really good

example. You have whatever your stocks

were doing earlier this month will have

a huge effect of what they're doing

today if you're investing. So having an

RNN model, a recurrent neural network

feeding into itself what was happening

previously allows it to take that model

and program in that whole series without

having to put in the whole a month at a

time of data. You can only put in one

day at a time. But if you keep them in

order, it will look back and say, "Oh,

this because of what happened yesterday,

I need some information from that and

I'm going to use that to help predict

today's." And so on and so on. We're

going to go back to our activation

functions. Remember I told you uh ReLU

was one of the most common functions

used. Uh so let's talk a little bit more

about ReLU and also softmax. Softmax is

an activation function that generates

the output between zero and one. It

divides each output such that the total

sum of the outputs is equal to one. It

is often used in the output layers.

Softmax L of the N equals E to L the N

over the absolute value of E to the L.

So what does this function mean? I mean

what is actually going on here? So we

have our uh output nodes and our output

nodes are giving us uh let's say they

gave us 1.2.9 and point4. As a human

being I look at that and I say well the

greatest value is 1.2. So whatever

category that is if you have three

different categories maybe you're not

just doing if it's a cat or it's a dog

or u oh let's say it's a cow. We had

cats and dogs earlier. Why the cats and

dogs are hanging out with a cow. I don't

know. But we have a value and it might

say 1.2 2 is a cat, 0.9 is the dog, and

point4 is a cow. Uh, for some reason, it

thinks that there's a chance of it being

any one of these three items, and that's

how it comes out of the output layer.

Well, as a human, I can look at 1.2 and

say this is definitely what it is. It's

definitely a cat or whatever it is. Uh,

maybe it's looking at different kinds of

cars might be a better whether it's a

car, truck, or a motorcycle. Maybe

that'd be a better example. Well, from a

computer standpoint, that might be a

little confusing because they're just

numbers waving at us. And so with the

soft max, we want all those numbers to

always add up to one. So when I add

three numbers together, I want the final

output to be one on there. And so it

goes through this formula changes each

of these numbers. In this case, it

changes them to 46.34

and 2. They all add up to one. And

that's a lot easier to register because

it's very set. It's a set output. It's

never going to be more than one. It's

never going to be less than zero. And so

you can see here that there's probably a

pretty high chance that it's the first

one. So you're as a human being, we have

no problem knowing that. But this output

can then also go into say another input.

So it might be an automated car that's

picking up images and it says that image

in front of us is probably a big truck.

We should deal with it like it's a big

truck. It's probably not a motorcycle.

Um or whatever those categories are.

That's the softmax part of it. But now

we have the ru. Well, what where's the

ru coming from? Well, the ru is what's

generating the 1.2 and the 0.9 and the

point4. And so if you remember our relu

stands for rectified linear unit and is

the most widely used activation

function. We looked at a number of

different activation functions including

tangent h the step function. Remember I

said the step function is really used if

that's what your actual output is

because then you know it's a zero or

one. But the relu if you have that as

your output you now have a discrepancy

in there. And if that's going into

another neural network or another

process having that discrepancy is

really important. and it gives an output

of x if x is positive and zero

otherwise. So it says my x value is

going to be somewhere between zero or

one and then the uh usually unless it's

really uncertain the output's usually a

one or zero and then you have that

little piece of uncertainty there that

you can send forward to another network

or you can look at to know that there's

uncertainty involved and is often used

in the hidden layers. This is what's

coming out of the hidden layers into the

output layer usually or as we reference

the uh convolution neural network the

CNN you'd have to go to another video to

review the RLU is the most common used

for convolutional part of that network

has a bunch of little pieces that are

very simplified looking at all the

different images or different sections

of the map and the RLU works really good

for that like I said there's other

formulas used but that this is the most

common one and you'll see that in the

hidden layers going maybe between one

layer and the next layer. So just a

quick recap, we have our soft max, which

means that if you have uh numerous

categories, only one of them is going to

be picked, but you also want to have

some value attached to it, how well it

picked it, and you put that between 01.

So it's very uh standardized. So we have

our soft max. We looked at that. Let's

go back one. We looked at that here

where it transforms the numbers. And

then we have our ReLU function which

takes the information in the summation

and puts it between a zero and a one

where it's either clearly a zero or

depending on how confident our model is,

it'll go between the zero and one value.

What are hyperparameters? Oh, this is a

great interview question.

Hyperparameters. When you are doing

neural networks, this is what you're

playing with most of the time once

you've gotten the data formatted

correctly. A hyperparameter is a

parameter whose value is set before the

learning process begins. Determines how

a network is trained and the structure

of the network. This includes things

like the number of hidden units, how

many hidden layers are you going to have

and how many nodes in each layer.

Learning rate. Learning rate is usually

multiplied once you figured out the

error and how much you want to change

the weights. We talked about or I

mentioned it earlier just briefly. You

don't want to just make a huge change

otherwise you're going to have a biased

model. So you only take little

incremental changes and that's what the

learning rate is is those small

incremental changes. Epics, how many

times are you going to go through all

the data in your training set? So one

epic is one trip through all the data.

And there's a lot of other things

depending on which model you're working

with and which programming script you're

working with. Like the Python sklearn

package will have it slightly different

than say Google's TensorFlow package

which will be a little bit different

than the Spark machine learning package.

So these are just some examples of the

hyperparameters. And so you see in here

we have a nice image of our data coming

in and we train our model. Then we do a

comparison to see how good our model is.

And then we go back and we say, "Hey,

this this model's pretty good, but it's

biased." So then we send it back and we

change our hyperparameters to see if we

can get an unbiased model or we can have

a better prediction on it that matches

our data closer. What will happen if

learning rate is set too low or too

high? We have a nice couple graphs here.

We have one over here. It says a

learning rate set too low. And you can

see that it slowly works its way down

the curve. And on the right you can see

a learning rate set too high. It's just

bouncing back and forth. When your

learning rate is too low, that's what we

studied two slides ago. That's what the

learning rate was. Training of the model

will progress very slowly as we are

making very tiny updates to the weights.

We'll take many updates before reaching

the minimum point. So I just mentioned

epic going through all the data. You

might have to go through all the data a

thousand times instead of 500 times for

it to train. Learning rate too high

causes undesirable divergent behavior to

the loss function due to drastic updates

and weights. At times it may fail to

converge or even diverge. So if you have

your learning rate set too high and it's

training too quickly, maybe you'll get

lucky and it trains after one epic run,

but a lot of times it might never be

able to train because the weights are

changing too fast. They they flip back

and forth too easy. And you see down

here we've introduced uh two new terms

converge and diverge. Converge means

that our model has reached a point where

it's able to give a fairly good answer

for all the data we put in. All those

weights have adjusted and it's minimized

the error. Diverge means that the data

is so chaotic that it can never manage

to to train to that data. The data is

just too chaotic for it to train. So we

have two new words there. Converge and

diverge are important to know. Also what

is dropout and batch normalization?

Dropout is a technique of dropping out

hidden and visible units of a network

randomly to prevent overfitting of data.

It doubles the number of iterations

needed to converge the network. So here

we have our standard neural network and

then after applying dropout. Now it

doesn't mean we actually delete the

node. The node is still there and we're

still going to use that node. What it

means is that we're only going to work

with a few of the nodes. Um, a lot of

times I think the most common one right

now used is 20%. Uh, so you'll drop out

20% of the nodes when you do your

training, you reverse propagate your

data and then you'll randomly pick

another 20 nodes the next time you go

through an epic data training. So each

time you go through one epic, you will

randomly pick 20 of those nodes not to

not to mess with. And this allows for

less overfitting of the data. So by

randomly doing this you create some I

guess it just kind of pulls some nodes

off to the side and says we're going to

handle the data later on so we don't

overfit. Batch normalization is a

technique to improve the performance and

stability of neural network. The idea is

to normalize the inputs in every layer

so that they have mean output and

activation of zero and standard

deviation of one. This question covers a

lot of different things which is great.

It's a great uh interview question

because it pulls in that you have to

understand what the mean value is. So a

mean output activation of zero that

means our average activation is zero. So

when you normalize it remember usually

we're going between minus1 and one on a

lot of these. It's a very standard

setup. So you have to be very aware that

this is your mean output activation of

zero. And then we have our standard

deviation of one. So we want to keep our

error down to just a one value. The

benefits of this doing a batch

normalization is it provides

regularization. It trains faster, higher

learning rates and weights are easier to

initialize. What is the difference

between batch gradient descent and

stochastic gradient descent? Batch

gradient descent. Batch gradient

computes the gradient using the entire

data set. It takes time to converge

because the volume of data is huge and

weights update slowly. So you can look

at the batches. A lot of times if you're

using big data, batch the data in, but

you still go through a full epic. You

still go through all the data on there.

So bash gradient descent means you're

going to use it to fit all the data and

look for a convergence there. Stochastic

gradient descent. Stochastic gradient

computes the gradient using a single

sample. It converges much faster than

batch gradient because it updates weight

more frequently. Explain overfitting and

underfitting and how to combat them.

Overfitting happens when a model learns

the details and noise in the training

data to the degree that it adversely

impacts the execution of the model on

the new information. It is more likely

to occur with nonlinear models that have

more flexibility when learning a target

function. An example of this would be um

if you're looking at say cars and trucks

and motorcycles, it might only recognize

trucks that have a certain box-like

shape. It might not be able to notice a

flatbed truck unless it's only a

specific kind of flatbed truck or only

Ford trucks because that's what it saw

on the training set. This means that

your model performs great on your train

data and great on maybe a small test

amount of data, but when you go to use

it in the real world, it leaves out a

lot and start and is not very functional

outside of your small area, your slow

laboratory data coming in. Underfitting,

doing the opposite when you underfit

your data. Underfitting alludes to a

model that is neither well-trained on

training data nor can generalize to new

information. Usually happens when there

is less and improper data to train a

model. has a bur performance and

accuracy. So if you're using underfitted

data and you generate a model and you

distribute that in a commercial zone,

you'll have a lot of people unhappy with

you because it's not going to give them

very good answers. So we've explained

overfitting and underfitting. So now we

want to ask how to combat them.

Combating overfitting and underfitting,

resampling the data to estimate the

model accuracy, k-fold cross validation,

having a validation data set to evaluate

the model. So when we do the reampling,

we're randomly going to be picking out

data and we'll run it a few times to see

how that works depending on our random

data and how we sample the data to

generate our model and then we want to

go ahead and validate the data set by

having our training data and then

keeping some data on the side uh testing

data to validate it. How are weights

initialized in a network? Initializing

all weights to zero. All the weights are

set to zero. This makes your model

similar to a linear model. So if you

have linear data coming in, doing a

basic setup like that might work. All

the neurons in every layer perform the

same operation given the same output and

making the deep net useless. Right?

There's a key word. It's going to be

useless if you initialize everything to

zero. At that point be looking into some

other uh machine learning tools.

Initializing all weights randomly. Here

the weights are assigned randomly by

initializing them very close to zero. It

gives better accuracy to the model since

every neuron performs different

computations. And here we have the

weights are set randomly. We have our

input layer, the hidden layers and the

output layer. And W equals NP random

random N layer size L, layer size L

minus one. This is the most commonly

used is to randomly generate your

weights. What are the different layers

in CNN? Convolutional neural network.

First is the convolutional layer that

performs a convolutional operation. We

have our other video out if you want to

explore that more. and go into detail

exactly how the C the convolutional

layer works in the CNN as far as

creating a number of smaller uh picture

windows that go over the data. Uh the

second step is as a relu layer relu

brings nonlinearity to the network and

converts all the negative pixels to

zero. Output is rectified feature map.

So it goes into a mapping feature there.

Pooling layer pooling is a down sampling

operation that reduces the

dimensionality of the feature map. So we

have all our relu layer which is pulling

all these little maps out of our

convolutional layer. It's taking that

picture and little creating little tiny

neural networks to look at different

parts of the picture. Uh then we need to

pull it together and then finally the

fully connected layer. So we flatten our

pooling layer out and we have a fully

connected layer recognizes and

classifies the objects in the image. And

that's actually your forward propagation

reverse propagation training model

usually. I mean there's a number of

different models out there of course.

What is pooling in CNN and how does it

work? Pooling used to reduce the spatial

dimensions of a CNN performs down

sampling operation to reduce the

dimensionality. Creates a pulled feature

map by sliding a filter matrix over the

input matrix. I mentioned that briefly

on the previous slide. Um it's important

to know that you have if you see here

they have a rectified feature map. And

so each one of those colors like the

yellow color that might be one of the a

smaller little neural network using the

ReLU. You'll look at it'll just kind of

um go over the main picture and look at

all the different areas on the main

picture. So you might step one 2 3 four

spaces. Um and then you have another one

that's also looking at features and it

has a 2785. Each one of those is a map.

So it might be the first one might be a

map looking for cat ears and the second

one looking for human eyes. When it does

this, you then have this rectified

feature map looking at these different

features and the max pooling with a 2x

two filters and a stride of two. Stride

means instead of skipping every pixel,

you're going to go every two pixels. You

take the maximum values and you can see

over here we look at a pulled feature

map. One of the features says, hey, I

had a max value of eight. So somewhere

in here we saw a human eye labeled as

eight. Pretty high label. And maybe

seven was a human hand and maybe four

was cat whiskers or something that we

thought might be cat whiskers. Four is

kind of a low number in this particular

case compared to the other ones. So you

have your full pool feature map. You can

see the process here is we have our

stepping, we look for the max value and

then we create a poolled feature map of

the maxed values. How does a LSTM

network work? That's long shortterm

memory. So the first thing to know is

that an LSTMs are a special kind of

recurrent neural network capable of

learning long-term dependencies.

remembering information for long periods

of time is their default behavior. We

did look at the RNN briefly talked about

how the hidden layer feeds back into

itself. With the LSTM has a much more

complicated feedback and you can see

here we have the hidden layer of T minus

one and the hidden layer that's what the

H stands for hidden layer of T and the

formulas going in. As we can see here we

have the hidden layers we have T minus

one and then H of T where T stands for

time. So this is a series remember

working with series and we want to

remember the past and you can see you

have your ex your input of t and that

might be a frame in a video as a frame

comes in they usually use in this one

the tangent h activation formula but you

also see that it goes through a couple

other formulas the omega formula and so

when it combines these that then goes

into the next layer your next hidden

layer that then goes into the data

that's submitted to the next input so

you have your x of t + one. So when you

have that coming in, then you have your

H value that's coming forward from the

last process. And depending on how many

of these omega structures you put in

there depends on how long-term the

memory gets. So it's important to

remember this is more for your long-term

recurrent neural networks. The three

steps in an LSTM, step one decides what

to forget and what to remember. Step

two, selectively update cell state

values. So based on what we want to

remember and forget, we want to update

those cell values and then decides what

part of the current state make it to the

output. So now we have to also have an

output on there. What are vanishing and

exploding gradients? This is a great

question that affects all our neural

networks. While training an RNN, your

slope can become either too small or too

large and this makes the training

difficult. When the slope is too small,

the problem is known as vanishing

gradient. So our slope, we have our

change in x and our change in y. When

the slope decreases gradually to a very

small value, sometimes negative, and

makes training difficult. When the slope

tends to grow exponentially instead of

decaying, this problem is called

exploding gradient. The slope grows

exponentially. You can see a nice graph

of that here. Issues in gradient

problem, long training time, poor

performance, and low accuracy. What is

the difference between epic, batch, and

iteration in deep learning? Epic. An

epic represents one iteration over the

entire data set. So that's everything

you're going to go ahead and put into

that training model. Batch. We cannot

pass the entire data set into the neural

network at once. So we divide the data

set into a number of batches. And then

iteration. If we have 10,000 images as

data and a batch size of 200, then the

epic should run 10,000 times over 200.

So that means we have our total number

over the 200 equals 50 iterations. So in

each epic we're running over all the

data set, we're going to have 50

iterations. And each of those iterations

includes a batch of 200 images in this

case. Why TensorFlow is the most

preferred library in deep learning? Uh

well, first TensorFlow provides both C++

and Python APIs that makes it easier to

work on. Has a faster compilation time

than other deep learning libraries like

KAS and torch. TensorFlow supports both

CPUs and GPUs computing devices. So

right now TensorFlow is at the top of

the market because it's so easy to use

for both programmer side and for

hardware side and for the speed of

getting something up and running. What

do you mean by tensor in TensorFlow?

Tensor is a mathematical object

represented as arrays of higher

dimensions. These arrays of data with

different dimensions and ranks that are

fed as input to the neural network are

called tensors. And you can see here we

have a tensor of dimensions five, four.

So it's a two-dimensional tensor coming

in. Um, you can look at an image like

this that each one of those pixels is a

different value if it's a black and

white. So, it might be zero and ones and

then each one represents a black and

white image. In a color photo, you might

um either find a different value system

or you might have a tensor value that

has the xy coordinates as we see here

plus the colors. So, you might have

three more different dimensions for the

three different images, the red, the

blue, and the yellow coming in. And even

as you go from one layer or one tensor

to the next, these layers might change.

We might flatten them, might bring in

numerous. In the case of the convergence

neural network, we have all those

smaller different mappings of features

that come in. So each one of those

layers coming through is a tensor. If it

has multiple dimensions coming in and

weights attached to it, what are the

programming elements in TensorFlow?

Well, we have our constants. Constants

are parameters whose value does not

change. To define a constant, we use

tf.constant command. Example, A equals

TF.Constant 2.0 TF float 32. So it's a

tensor float value of 32. B equals TF

constant 3.0. Print AB. If we did a

print of AB, we'd have um TF.stant and

then of course uh B is that instance of

it. Variables. Variables allow us to add

new trainable parameters to graph. To

define a variable, we use TF.variable

command and initialize them before

running the graph in session. Example W

equals TF variable.3 DT type TF float 32

or B equals a TF variable minus 3, D

type float 32. Placeholders.

Placeholders allow us to feed data to a

TensorFlow model from outside a model.

It permits a value to be assigned later.

To define a placeholder, we use TF

placeholder command. Example A equals TF

placeholder b= a * 2 with the TF session

as SESS result equals session run B,

feed dictionary equals A3.0. 0 print

result. Uh so we have a nice example

there of a placeholder session. A

session is run to evaluate the nodes.

This is called as the tensorflow

runtime. So for example, you have a= tf

constant 2.0 b= tf constant 4.0 c= a

plus b. And at this point you'd go ahead

and create a session equals tf session.

And then you could evaluate the tensor C

print session run C. That would input C

as an input into your session. What do

you understand by a computational graph?

Everything in TensorFlow is based on

creating a computational graph. It has a

network of nodes where each node

performs an operation. Nodes represent

mathematical operation and edges

represent tensors. Since data flows in a

form of a graph, it is also called a

data flow graph. And we have a nice

visual of this graph or graphic image of

a computational graph. And you can see

here we have our input nodes, our add

multiply nodes and our multiply node at

the end. And then we have the edges

where the data flows. So we have from A

going to C, A going to D. You can see we

have a two flowing, a four flowing.

Explain generative adversarial network

along with an example. Suppose there is

a wine shop that purchases wine from

dealers which they will resell later. So

we have our dealer going to the wine,

our shop owner that then sells it for a

profit. But there are some malfactor

dealers who sell fake wine. In this

case, the shop owner should be able to

distinguish between fake and authentic

wine. The forger will try to different

techniques to sell fake wine and make

sure certain techniques go past the shop

owner's check. So, here's our forger

fake wine shop owner. The shop owner

would probably get some feedback from

the wine experts that some of the wine

is not original. The owner would have to

improve how he determines whether a wine

is fake or authentic. Goal of forger to

create wines that are indistinguishable

from the authentic ones. Goal of shop

owner to accurately tell if the wine is

real or not. There are two main

components of generative adversarial

network. And we refer to as a noise

vector coming in where we have our

forger who's going to generate fake wine

and then we have our real authentic wine

and of course our shop owner who has to

figure out whether it's real or fake.

The generator is a CNN that keeps

producing images that are closer in

appearance to the real images while the

discriminator tries to determine the

difference between real and fake images.

The ultimate aim is to make the

discriminator learn to identify real and

fake images. What is an autoenccoder?

The network is trained to reconstruct

its inputs. It is a neural network that

has three layers. Here the input neurons

are equal to the output neuron. The

network's target outside is same as the

input. It uses dimensionality reduction

to restructure the input. Input image

comes in. We have our Latin space

representation and then it goes back out

reconstructing the image. It works by

compressing the input to a Latin space

representation and then reconstructing

the output from this representation.

What is bagging and boosting? Bagging

and boosting are ensemble techniques

where the idea is to train multiple

models using the same learning algorithm

and then take a call. So we have in here

where we're bagging. We take a data set

and we split it. We're going to have our

training data and our test data. Very

standard thing to do. Then we're going

to randomly select data into the bags

and train your model separately. So we

might have bag one, model one, bag two,

model two, bag three, model 3, and so

on. In boosting the emphasis is to

select the data points which give wrong

output in order to improve the accuracy.

So in boosting we have our data set

again we split it to test data and train

data and we'll take a bag one and we'll

train the model. Data points with wrong

predictions then go into bag two and we

then train that model and repeat.

machine learning is which is a subset of

artificial intelligence, right? That's

uh basically

um machines learning from data

in order to uh make decisions

essentially. Um so this was a big

departure from the rules-based systems

at the time, right? That were explicitly

programmed to make decisions. So just

think of an example like a really big

kind of if this then that then that then

that and and else if this this this

right so bunch of rules that had to be

pre-programmed in order to um come out

with some final answer. Uh with machine

learning it's the exact opposite of

that. we're actually training something

from examples from existing data um in

order to predict something or um

make some type of decision. Uh and so

we're going to learn about the various

ways we can do machine learning. But if

you guys remember we um talked about

some of this like the differences and

the uh basically rules-based approaches

to learning from data approach. Um and

in included in that is going to be uh

complex unstructured data. So things

like images, text, audio. What handles

those really well is uh deep learning

which we will get to in the course after

this. But uh those are certainly in

there as learning from data even complex

data.

So we had this picture uh and I think

this is kind of around where we left off

last time was uh just distinguishing

between those three terms. We see

artificial intelligence, deep learning

and machine learning kind of used

interchangeably, but this is really how

they fit in. Artificial intelligence is

kind of a broad anything mimicking human

intelligence. Um which doesn't have to

be learning from data, but uh machine

learning is part of that. And then um

one way to accomplish machine learning

is to use neural nets which is the focus

of uh deep learning. Um and so deep

learning has been has found a lot of

success especially recently with uh

those complex data types like images,

speech, text, right? So deep learning

used all over the place. Even in um

modern like generative AI, we see deep

learning used quite a bit. Um it really

anything that's using neural nets is uh

going to be deep learning.

Um and again we'll focus on that later

but we're going to be mainly focused on

machine learning for this course.

Primarily machine learning that does not

use neural networks. Okay. So just

models that are not necessarily neural

networks

be our focus.

So in machine learning we had an example

of a game uh essentially um learning

what decisions to make uh based on the

uh kind of current um state of the

board. This could be a um you know

machine learning example that uh learns

from many previous examples. So a lot of

data around these games are used to

train these um kind of robots that can

play these games and play them at a very

high level. Um so there's been a lot of

successes actually in machine learning

and deep learning um around

uh playing games like chess or go

um using machine learning algorithms. So

pretty cool.

All right, so I think this is where we

ended. Last time we said there's a bunch

of different use cases for machine

learning. So um recommendation system is

going to be a big one and we will

actually study that uh in one of our

final lessons of this course. Um chat

bots like generative AI doing sentiment

analysis chat bots we'll study later but

those are certainly an application of

learning from data in order to uh

generate responses to text prompts

right. Um spam filtering that's a good

example like classifying an email as

spam or not spam. Um that that gets

trained from examples and uh learning

from data such as previous emails. Um

social media posts analysis is another

kind of text data um use case but you uh

can do a lot with that text like you can

predict the sentiment um you can predict

uh the category of what what the post is

talking about um those kind of things

all can be done with machine learning

>> and many other use cases not on this

list that we will uh cover

>> you know as we as we go further.

Okay, so this is where we kind of left

off. Um, so what's doing all the hard

work here is

>> uh machine learning algorithms. So these

are things that will um these are things

that will learn from the data. So they

are uh they they are basically um

algorithms or sets of rules that uh or

mathematical rules I should say not

formal rules like in the in the sense of

a rule system but mathematical um

formulas and mathematical uh rules

essentially that help us learn from the

data. So they correlate the data to some

type of outcome. So some type of

prediction uh whether that's going to be

as we will see whether that could be

like a number like we're predicting a

price or demand or sales

um or it could be a category like is

this transaction fraud or not fraud or

what's the probability that this is

fraud um so we have different kinds of

predictions we can make with machine

learning

um but uh we will study the kind of the

differences of those coming up. Um, but

machine learning algorithms are really

what power they're kind of the models,

right? They're the models that help

power uh machine learning to actually

learn from data.

So, we're going to spend a lot of time

in this course studying those algorithms

like the different models that we can

build and what their differences are,

what their strengths are, what their

weaknesses are. We'll we'll learn a lot

about those.

Okay. So I guess you can imagine like

everything is so data dependent, right?

Um we're learning from data. So uh it

makes sense that the quality of data

really really matters here in

determining how strong the model can be.

Um so you see this graph here charting

kind of the um high quality data um

versus just uh any old data but a decent

enough quantity of it. Um you can see

that performance and the performance is

measured by some evaluation metric. Um,

so think of it as uh something like an

accuracy. Like if we were predicting

fraud or not fraud, how accurate can our

model get at actually detecting fraud,

um, it gets better and better and better

the graph shows that the higher quality

of data that we have. So there's kind of

that there's a there's a saying in

machine learning um, called garbage in

garbage out. What that means is if you

have poor data, even the best model in

the world, poor data is not going to

result in having a good model that can

be accurate and perform well. Um, so it

needs to be high quality, meaning um

there needs to be a decent amount of it

and it needs to be labeled appropriately

as we will will talk about

um and it needs to not have any, you

know, significant outliers. it needs to

be clean, not have those missing values,

all of those things. Um, you can you

have a good chance at deriving good

predictions from higher quality data

as this kind of shows.

Okay.

So, one thing we're going to learn um as

we go along is

quantity matters as well. So, not only

quality, but a decent amount of it. And

um we're going to learn those kind of

rules of thumb like how much data do I

need for certain algorithms. Um one

thing that we will see is that uh the

the basic machine learning models that

we'll study don't need as much as a

neural network would. It you know neural

networks are going to require a lot more

um than a basic machine learning model

learning model. So uh that's something

we will see as we go along. But uh this

is something we'll talk about and

discuss with each model that we study is

kind of how much data do we actually

need to produce a high quality model.

Okay, any questions uh so far?

Okay, let's talk about the different

types of machine learning that we're

going to discuss. Pime, there's going to

be two primary ones that we will study

in this course and then a couple others

that'll be a little bit more advanced

that we won't get to but worth knowing

about. Um, so there's going to be four

total that we'll study or talk about and

they'll be on this list here, which is

um supervised learning and unsupervised

learning. Now, I'd say the majority of

our focus will probably be on supervised

learning, and we'll talk about what that

means, but we'll also cover unsupervised

learning as well. And so, we'll look at

the most popular techniques in each of

these types of machine learning.

Um,

and then we'll talk about these two, but

not really study them because they're

more advanced topics um that that will

be beyond the scope of what we'll do.

But uh these are going to be um

different styles of machine learning

that are going to be characterized by um

what kinds of predictions they make,

what kind of data they need and require.

Um and uh what kind of outcomes they're

actually producing. Um, so let's let's

get into each of these, but uh the the

one that we'll probably spend the

majority of our time on is going to be

supervised learning, but we will study

unsupervised learning as well. We'll

study both and we're going to talk about

we're going to define both of those um

coming up. And again, these will be a

little bit more advanced topics that we

won't spend too much time on.

Um, but but we'll discuss their

relevancy in machine learning um and

give a good definition to it.

Okay.

All right. Let's start with supervised

learning. Now this is going to be uh a

term that really refers to

using examples. So using labeled

examples. So here we say labeled data to

help our model train. In other words,

help our model be able to predict guided

by specific input output pairs. So

supervised really refers to the fact

that we have answers. We have examples,

we have answers with those and we use

that collection of data to build our

model off of so that we can predict

um those kinds of things like a price,

like a category, like a spam not spam.

in this in this slide like we would be

predicting if this shape is a square, a

triangle or a circle.

Um but but when we build a model for

that, we have data that has an answer

attached to it. Right? We've talked

about this before a little bit with

labels. So there's a guide there that

can guide us towards building our model.

there's an actual every every example

has an answer and that answer is really

critical to help build our model off of.

So, um that's it's almost like you have

um a you have a bunch of exercises

in let's say like a math textbook. You

have a bunch of exercises and you have

the answers and that way you can kind of

check your work. You think about model

training um that is the really a lot of

that process of model training as we are

going to discover is um basically

checking our work against these answers

in our data in our training data.

Okay. So supervised learning is any type

of machine learning that involves

learning from labeled data in order to

predict outcomes. Okay, predict outcomes

like now the the outcomes can be

numerical. They can be like a price,

temperature, demand, sales, revenue.

They can be numerical, but they can also

be categorical. So they can be like

spam, not spam, fraud, not fraud,

cancer, not cancer. Um, dog, cat,

giraffe, those kind of categories. Um,

we could predict those. It's some type

of outcome. Okay, some type of outcome.

The key is we're using labeled examples

to guide our model building. That's why

it's called supervised learning.

So we know in our data we know what the

inputs are. Of course, those are going

to be think of the inputs as like all of

our columns and then we have a special

label column that represents the output

we're trying to predict. So if you think

about that housing price data, the label

could be the price. And that's something

we would build a model to predict, but

we have answers for all of our examples

in our rows. We have answers to help

guide our model building.

They help tweak our model because we

know the answer ahead of time. So

they're they're really good examples to

build our model off of.

Okay. So that's that's supervised

learning.

Uh in this example is circle not in the

prediction because it's not part of the

test data even though it's in the

labeled data.

Um no it just not necessarily. It just

means that like we learn against all of

these examples that have these answers

and then when we observe new examples um

we can try to predict what those would

be based on what we've seen before. So I

if there was a you know it's just a

coincidence we only have two two

examples in our test data like we could

have a circle here in which case we

would predict circle

that's fine or at least we would hope

our model would predict circle right

that's what we're hoping may or may not

get it right

um but it's it's only not there because

we only like we're just assuming that we

only have two examples we're testing

against but in reality we would probably

do a lot more than two.

It's just it's just a coincidence

really.

In reality, we would test against a lot

more data. And we're actually going to

see why we would do that. Like why would

we train our model and then kind of use

additional data um to to evaluate it?

It's actually really important that we

do that step to get a sense of how good

our model is before we take it out in

the real world. So if we apply our model

that we build on our label data to

um this kind of set of test data that we

haven't been exposed to before. It helps

give us a sense of how good is our

model. So it's tested is usually used

for evaluation.

So that's something that's something

we'll study.

How do we train? Uh it depends on the

model. Um so training will be a sense uh

will be an algorithm that will um

basically update the model according to

the data. These labeled examples. Um

every model is going to be different in

exactly how it trains. So we're going to

we're going to talk about that when we

get to the individual models that we'll

study.

But uh loosely speaking, they're going

to use the data to adjust itself. Like

imagine adjust like tuning a bunch of

knobs. Um, like the best example I can

give you is we I think I did this one

last week where you have kind of a

function

that predicts the price and let's say it

has

um weights like weight one with feature

one, weight two with feature two,

weight three with feature three. So

imagine we had three input features and

we we built an answer according to that.

Essentially what we would do to train

the model is adjust these

um in order to get this correct based on

our our labeled examples.

Okay.

So that's something we're going to learn

about coming up shortly when we when we

actually dive into model. Every model is

going to be slightly different in how it

trains, but at a high level it's going

to use the training data with those

examples, right? the labeled examples to

help guide the formula essentially to

adjust to generate the proper kind of

model here.

The these things are going to be

adjusted according to the data

in order to produce the correct output.

So think about these as knobs that will

turn.

Okay.

Uh which type of machine learning is

used? Uh probably supervised um which is

what we're talking about now. So

probably supervised because most people

want to

um build some type of model to predict

something.

Uh so yeah, I'd say I'd say supervise.

Yes, we're are we are definitely going

to learn how to train. Yeah, we'll see.

We'll do the code. Um, I'll tell you

about how it's done. Yeah, we're

definitely going to learn it. But what I

was saying is it's kind of on a model

bymodel basis.

So, I want to wait till we get into the

individual models, then we'll talk about

how they're trained.

But yeah, we'll we'll learn how to do

that.

But yeah, supervisor is used all over

the place. Even even for uh generative

models, they use supervised learning

because um like an LLM

is going to use labeled examples in

order to train, right? In order to train

how to generate responses according to

prompts. Um it needs to learn against a

lot of text examples.

So that supervised learning is what um

results in that model,

right? Learning from those labeled

examples.

Okay.

It is yeah image image uh a lot of um

yeah a lot of image processing is

supervised like object detection. So the

YOLO model is an object detection model.

Yes. Um because it has to be trained

right. It has to be trained on uh it has

to be trained on images

with labels such as this is what object

is in this image. This is the box around

the object.

Um yes. So if if it's if it ever uses

label data to train and build the model,

it is supervised. So YOLO is definitely

supervised and we actually we will we

will cover the YOLO model later on in

our deep learning course. We talk about

object detection.

So we'll we'll study that.

But yeah, it's supervised

Okay. So on the slide we have some

common supervised learning algorithms

that are we will study. So all of these

we will study and understand what they

do and how they work but just giving you

some to name them. linear regression is

kind of the one I just drew out which is

the um this is the prototypical like

easiest to understand model that is kind

of the um exactly like this where we

have a weight times a feature

um a weight times a feature and then a

weight times a feature

and on and on and on. You can have as

many as you want.

um that is a linear regression. And so

that is um that's a supervised model

because we need this value here and we

need all of our inputs in order to um

actually train this model and generate

all those weights

um that that is uh that uses um labeled

examples to help tune all those knobs.

Um same with all these other models. So,

we're going to talk about decision

trees. We're going to talk about

logistic regression and and SVMs, which

are support vector machines. We'll talk

about all of those, but they're all

examples of supervised uh supervised

learning.

Okay, we'll talk about all of these.

They're all supervised because they all

require labeled examples in order to

train them and and then subsequently use

them. Okay.

Okay. So what are some use case

examples? So for for instance in uh

supervised learning we may be predicting

temperature based on yearly temperature

trends. So we would have that yearly

data as our um as our labeled examples

and those would supervise the learning

of a model that predicts temperature.

Um, same thing with predicting crop

yield based on um, seasonal crop quality

changes. So maybe we have a bunch of

features relating to crop quality. We

could predict crop yield. Um, we would

just need historical examples with those

labels, right? What the crop yield is

for each time period. Let's say we would

just need those uh, supervised examples

and we could easily build a model off of

it.

Um

uh this this last one sorting waste

based on known waste items and their

corresponding waste types. Um that's

kind of like spam. It's like filtering

basically like a spam filtering. Um so

think of it like the the shapes example.

We sorting things into squares, circles,

triangles. Um, same kind of idea here

where we have a bunch of examples on

what those um what those waste items

should uh should belong to, like what

wastist bins they would go to, for

example. Um, and those could be labeled

and therefore then we could um

understand what category of waste they

belong to.

Um, same thing with spam. Something is

fraud or not fraud. spam or not spam,

cancer or not cancer. All of those are

going to be supervised learning examples

because they're going to require in

order to train them, they're going to

require data that has those labels.

Okay? So, anything that has labels is

going to be supervised learning. So

again, this is where we will spend

probably the the majority of our time is

doing supervised learning problems, ones

that we have labeled data. We're

building a model and we're going to

predict those those uh labels

essentially.

Okay, before we go to unsupervised, any

questions about uh supervised.

Okay.

All right. So supervised requires labels

in order to have an example to go off of

to build your model. And that's because

you're predicting those kind of outcomes

like spam or not spam, cancer or not

cancer. Now unsupervised learning is

completely different. It's the opposite.

So unsupervised learning is where we do

not use labels whatsoever. So we're not

using any labels at all. So it's it it

can be completely unlabeled or even if

it's labeled, we're not using labels in

any way. But um we primarily would say

it's unlabeled data. We have no guidance

because we're not using the labels in

any way. We have no guidance to um

predict anything, but that's because

we're not really predicting anything in

unsupervised learning. Generally, what

we're doing is looking for some

structure or pattern.

Okay, with unsupervised learning, we're

looking for some structure or pattern.

So, um, one type of example that's very

very popular is going to be this second

one, which is, um, identification

identification of user groups based on

similarities or commonalities. Now, this

is going to be a problem basically known

as clustering

and it's a problem we will study quite a

bit. there's going to turn out to be

lots of different algorithms that can

accomplish clustering. So what

clustering attempts to do is basically

say um we have data that's like this and

then data over here and then data over

here. Let's just group these together.

So like this should be one group, this

should be one group and this should be

one group. And we can find those

structures and say okay this is group

one, this is group two and this is group

three.

one, two, three. And we can basically

build what we would call clusters of

data um based on how close together the

points are kind of located in these kind

of cluster zones like these boxes I've

drawn.

Okay. Now, that doesn't require any

label to do which is really fascinating.

So, unsupervised, you don't need any

label at all to accomplish the

algorithm. Um so, clustering is one good

example.

um finding outliers or anomalies is

another. So we don't necessarily have

any label of what is an outlier or what

is an anomaly. We are deriving that from

the features alone. There's no guidance.

There's no label um to doing like

outlier detection or anomaly detection.

Okay, so that's another good example.

One that's not listed on here um but is

also really important that we will study

is something known as dimensionality

reduction.

So dim reduction and what that what this

focuses on is basically compressing the

data set a bit. So we take our data and

basically compress it um so that but we

do it in such a way that we retain as

much information as we can. This is a

very like smart compression and what it

does is it lowers the dimension. Um

dimension think of the dimension as like

number of columns.

Number of columns.

So imagine we had 100 columns in a data

frame. What we could do is actually

reduce that down to 10. So like 10% of

that. So we reduce it down to 10. And um

but those 10 are it's not like we

chopped out um 90 other columns. We um

smartly kind of compressed all that

information into these 10 new columns um

that are compressed versions of the

hundred that we used to have. Um so

dimensionality reduction is is another

unsupervised technique. It requires no

guidance, no label to do, but is um a

really useful technique to reduce the

size of your data if you're doing things

with it. Um so this is another one that

we will we'll study how to do it and

basically more details behind it, what

the algorithms are.

Um we'll so probably those two in

unsupervised will spend the most amount

of time on clustering and dimensionality

reduction.

uh and supervised if some data is

present but we didn't label it means in

example we had circle triangle square in

the training data we add pentagon

but we didn't label that in that case

uh yeah so every um in supervised

learning, every row, think about it as

like every row in our data frame needs

to have a label

uh associated to it. It needs to have a

a column that represents the label.

So if we've never seen Pentagon before,

I can't use that as a label.

So, it has to the Pentagon has to exist

in the data if I'm going to be able to

predict it,

right? So, I can't predict, right? If

we've never seen it before, we have no

examples to go off. We have no guidance.

So, how could we predict that?

Right? We can't predict it.

if it's if it's in there. So if if we

have labels of Pentagon, let's say, then

yeah, we could predict Pentagon. We

could

remove. Remove what?

We wouldn't if it was talking about the

Pentagon, we wouldn't remove that. No,

let me go back to that page. We wouldn't

remove it. Um, it's just if it's not in

our labels, we're not going to be able

to predict it. So, Pentagon's a good

example here. Uh, Pentagon is not one of

our labels. So, it currently is not in

our data set as one of the labels. We

only have data that's either a triangle,

circle, or a square. We don't have

pentagon. So, I would never be able to

predict pentagon. I'll never be able to

do that if I haven't seen examples of it

before.

Okay. But let's say we had that in

there.

So we had Pentagon.

So if we had Pentagon, um we could have

an example of it in our labels

and then yeah, we it could be then we

could predict it.

Yeah. Yeah. The the don't get worried.

Don't worry about the test data. So the

test data is just saying here's a new

here's a shape. What is it? Okay, that's

a square. Here's a shape. What is it?

Okay, that's a triangle. And we could

have as many of those examples as we

want in our test data. So we could have

a circle and say, okay, what's this?

Should be circle,

right? The test data can be whatever it

whatever it wants. But yeah, if if we've

never seen Pentagon before, we're never

going to be able to predict it.

These are the the label data and labels

are basically the talking about the same

thing. The labels just mean what are the

categories that are present in our data.

So in this data we only have three

labels that are present.

So the labels is are relative to our

label data, right? It's saying

what labels,

excuse me, what labels uh do we have

in our data and we only have those three

circle, triangle, square. So so Pentagon

would not be part of those labels. We

couldn't predict it.

No. So unsupervised is not going to make

a prediction. That's the big difference

with unsupervised. They're not going to

make a prediction like this. Um so

unsupervised is not going to make a

prediction. It's going to do something

different like um basically say like

these guys are similar, these are

similar, these are similar, this is a

cluster, this is a cluster, this is a

cluster. It's not going to make a

prediction. That's what supervised

learning does.

Clustering, yes, which is unsupervised,

yes, clustering does not require any

labels. Unsupervised just means we don't

have any labels. We don't require any

labels.

So the other thing unsupervised might do

is it might say

and again without the labels it might

say that this is an outlier.

it might say that this guy is an outlier

because there's only there's only one of

those and they're not like the other. So

that that's something that um that's

something that uh unsupervised could do.

Um it it yeah and no. It kind of labels

a cluster in the sense that um it would

basically assign a number to it like

this is cluster one, this is cluster

two, this is cluster three.

It'll assign a number to it, but it's

not a very meaningful it doesn't assign

like a prediction label in the in the

traditional sense of a label.

It does provide like a numerical index

for the cluster to because what we want

to know is like okay this guy has the

cluster of one. This guy belongs to

cluster one. This guy belongs to cluster

one. This guy belongs to cluster two.

This guy belongs to cluster two. Does

that make sense? So there needs to be

some like index of what cluster you

belong to.

So it's kind of like a label but not in

the traditional like prediction sense.

Very

good. So again, unsupervised, no labels.

You're doing things like

identifying clusters,

um identifying outliers, doing

dimensionality reduction. These are all

like structure and pattern oriented

things. They're not predictions of a

label. Okay? They're not which is what

we would see in supervised learning.

Okay. So an example would be that we

take we put in the data um we can group

together uh data such as images into

categories based on similarities um

which would be like those clusters. So

there's no these would be groups that we

don't have any label on ahead of time

like we don't have we don't say that

this image should belong to this this

image should belong to this we derive

that from the characteristics of the

data. Um so think like a good example is

um customer groups. So we would identify

customers based on like okay do they

have similar spending levels? How many

days do they go shopping in a week? How

much money do they spend? And we can

kind of group together customers based

on similar qualities.

Clustering will find those groups that

should exist.

um it will discover those groups based

on um the similarities in the data, but

there's no labels that that say like

this person should be in this group,

this person should be in this ahead of

time. There's no labels of that. It gets

derived during the algorithm. It's

unsupervised,

right? There's no unsupervised really

literally means no guidance. There's no

guidance to doing it. We just derive

that from the structure of the data

which is the similarities.

Okay.

All right. So,

a couple more for you. So we had um

supervised which uses the labels. We

have unsupervised which uses no labels

looking for structure. And then we have

something that's kind of in between

which is um what is known as

semiupervised learning. And this is

where you use a combination of a little

bit of label data, but most of your data

is actually unlabeled data. Um, and you

try to get some use out of that label

data in order to um build a model out of

it. And so uh it uses the um it uses

that label data to um generally provide

some guidance on usually what happens

with semi-supervised learning is you use

your label data to kind of predict what

the label should be for the unlabelled

data and then you can go from there. So

you can create artificial labels on this

unlabeled data and then you can use all

of it once it's all been labeled kind of

like a supervised learning uh approach.

So but but this is semi-supervised

basically refers to the fact that you

start out with most of your data not

being labeled but you do have some

labeled examples and what you can do is

basically extrapolate those labels into

the unlabeled data set and then provide

some artificial labels and then now

everything has a label you can do

supervised learning.

Okay. So, it falls kind of between um

supervised and and unsupervised.

Uh and there so this is this is kind of

rare. Most of the time you're not going

to do that. You're actually just going

to um prefer to just start with all

label data. That's usually the preferred

approach. Most of the time you'll

actually just be doing supervised

learning, not really semi-supervised

learning. So, it's pretty rare, but um

it it could like if Yeah, it could if

the if we had a lot of examples of

Pentagon and we wanted and so they were

unlabeled and then we tried to guess

what kind of shape they were um and

provide an artificial label uh and then

um then use that whole data set to build

a model off of then yeah it could it

could fall into this category. Okay.

They Oh, going back to the question,

they still use some kind of label data

like age, gender. They use uh that's

those aren't those aren't really labels.

That's the features. So, yeah, they

still use the core features of the data.

They just don't have any like labels in

the traditional sense of a label. Like

you should think of a label as something

we are trying to predict.

So whether that's a price, whether

that's like a category like spam, not

spam, cancer, not cancer, it's something

we'd be interested in kind of

predicting. And so um in our data, we

would have an answer for every row. We'd

have one of our columns would be like

the the result like the outcome answer

that we're trying to predict. That's the

label.

So in unsupervised, we don't have any of

the labels.

We do have just the regular features

like gender, age, income, square

footage,

bedrooms, bathrooms, all those things.

Okay.

So, we have semi-supervised that falls

in between supervised. Now, the reason

it falls between is be is because

there's a decent amount of data that's

unlabeled. In fact, a majority of it

unlabeled. But what we can do is try to

label it. We can try to take what we

know from our existing labels and

predict an artificial label and then use

all that data together in kind of a

supervised fashion for a model down the

road.

So that's kind of what this picture uh

says is we can try to take um you know

maybe we try to infer some labels based

on we have some some labelled data here.

We have most of our data is unlabeled

and we try to supply some labels to it.

Um like maybe we have a babies category

of teens, a tween, uh you know youth and

um adults. Um and then we try so we we

take our our labels and we try to

extrapolate those into artificial labels

for this unlabelled data so that we can

use it now because then everything has a

label at this point and then we can just

go ahead and do supervised learning from

there.

So we can do supervised from there. What

we would prefer to do and what we'll do

in this course

um is just start with supervised. We'll

just start with the labels. We won't try

to derive artificial labels usually.

We'll just start with labels.

So one example in the real world is

something like Google photos which um

whenever you take a picture it can

provide uh uh labels based on previous

uh images in your library. So it can it

can produce tags or um labels on those.

Uh generally when you take that picture

it's kind of unlabeled unless you go in

and specifically provide some tags and

some labels. But um if you don't do that

it can still it can still uh make it can

artificially create one of those based

on the other label data that you already

have.

So that's um

that's an example.

Okay.

All right. Last one in terms of machine

learning. So we have supervised, we have

unsupervised.

Uh then we had semi-supervised which is

somewhere in between a mixture of having

some unlabelled data and label data. Um

now we're going to talk about

reinforcement learning which is

completely different. Um it's it's

completely different than the other

three. It's a type of machine learning

where we uh basically learn from

interaction with the environment. And

you might ask what are we learning? We

are learning what actions to take in the

environment. Um and the way we do that

is by reinforcing

positive actions that lead to a a

reward. Um, so that's where the word

reinforcement comes from is we we

basically uh imagine like a child

that's, you know, learning from trial

and error. Like they're trying to crawl,

they're trying to walk and they keep

falling down. um eventually they learn

how to do it through trial and error and

they might get a reward or they might um

reinforce some of those positive

movements that lead them to walk or

crawl um or they might learn from the

penalties, right? They might learn from

uh some type of feedback. So they might

learn from falling down like, "Oh, that

hurts. I should uh support myself a

little bit better, right?" Or be a

little more coordinated. Um

and so they they learn from those

actions and their interaction with the

environment. Um

uh so this is a complex um algorithm

essentially uh it's it deals a lot with

um again taking actions. Usually when

you take an action something changes in

the environment um then you kind of

observe some type of feedback. So, think

about like a a board game where you're

trying to figure out what move you

should make or another good example is

like with a robot um trying to navigate

a maze. So, like what route should it

take? Should it move forward? Should it

move backward? Should it move left or

right? Those are different actions it

can take. Also, like a self-driving car,

should it should it turn? Should it

speed up? Should it slow down? Those are

all good examples of things that have

been trained from reinforcement

learning.

Uh yeah. So real world examples would be

like in a board game uh a a reward would

be like if you win the game. Um or if

you like capture a piece like in

checkers or chess, that's a reward. A

penalty would be like if you lose the

game or lose one of your pieces, that

could be a a penalty.

um in a board game or sorry in like a a

robot navigation task, it could get

rewards for um moving in the right

direction

um towards the exit or like when it like

let's say you wanted to train a robot on

how to open the door and navigate a

room. Um you would penalize it for

bumping into the wall.

Um you would give it a reward for moving

usually oh like oh the algorithms

themselves usually it's like a a step

function um it's usually it's like a

discrete function that kind of is based

on the state so the reward it could be

like um like depending on the let's

let's go back to the board game example

like the reward could be like or even

the maze let's say like a navigating the

maze like getting to this let's say this

was the exit

And this was the entrance.

Then if they make it to here, they get a

numerical like if they make it to the

exit, they get a numerical reward of

like plus 100, let's say. So it's just a

number. And then if they uh like if they

bump if they go into here, like let's

say this is kind of like a death trap or

like a pit, this this would be like a

minus 100. So it could be like discrete

numerical values could be the reward if

they're moving in the right direction.

Like let's say we want to encourage

going this way then we could give

smaller intermediate rewards like this

should be a plus like if you move

forward this is a plus five this is a

plus 10 this is a plus 15 if you're

moving in the wrong direction away from

the exit. Um that would be like a minus5

or a minus 10. Does that make sense? So

they're they're numerical in nature and

what you're trying to do is collect the

most reward. You're trying to get the

largest reward you can through trial and

error. So you you try this out many many

many times. You basically simulate

running through this maze many many

times. And what dictates it what

dictates like where I should go is based

on what I've observed in the past. It's

almost like you're a child remembering

like, okay, what move should I make from

this space? Like if I'm here, if I'm

here, which way should I go? Should I go

down? Should I go right? Should I go

left? You kind of know that from

experience.

Does that make sense? Based on the

reward that I've seen in the past, like

when I've moved down, I've gotten a

higher reward than moving left or right.

Does that make sense? So, yeah, it's

it's a numerical value

as a reward.

Yeah, that's a great question. Um, how

does it differentiate rewards based on

gain and loss? I chess. So it's it's a

very comp complicated uh answer but

essentially every so in the chess board

you can think of the board as like every

every um

space is a state.

So I could be in this state I could be

in this state and then it's not not only

is every every uh space but where all

the other pieces are. So there's lots of

states that are possible.

Um, so

the way there's a way to quantify

essentially what's the value of taking a

certain action like moving my piece

left, moving it right, moving it up or

down um given the rest of the state. So

you're you're right, it may be

beneficial to sacrifice. Um, but we

would learn that through experience that

okay, the best move in this situation is

to sacrifice.

We would we would have to learn that

through trial and error many many many

times which is to say like okay if I'm

in this current state of the world right

all these pieces are distributed in this

way the best move for me right now in

the long run

to get the most reward in the long run

is to actually sacrifice my piece and

move it right move it into like a bad

position theoretically but we know from

experience that's actually the most

long-term reward is from that position

like moving it right may be the best

for me. So what you learn is how to take

actions

and actions are usually like move right,

move left, move up, move down. You think

about like a self-driving car though,

that's going to be like slow down, speed

up, turn your wheel 10 degrees. Um those

kind of actions.

So the the short answer is it's there's

a calculation there that you learn what

the long-term value of every state is

every unique state

and then you're trying to basically say

what action should I take from that

state

given that current state of the

Okay.

And I really I really like reinforcement

learning. It's actually probably my

favorite field of machine learning.

Unfortunately, we won't be covering it

um in our main uh course. We have

offered uh electives around

reinforcement learning in the past. So,

um stay tuned. Maybe when we get to the

end of this program, uh we'll offer an

elective on it and if enough people sign

up for it, we'll we'll run it. But, um

we it's not part of our we don't really

cover reinforcement learning as part of

our main topics. It's it is an advanced

uh more advanced topic than than what

we'll cover, but um I I really enjoy it.

I find it very fascinating.

Okay. So, all of this is kind of um

illustrating what I was saying, which is

um you think of like uh the thing that's

interacting in the environment like the

robot or the car or the human moving a

chest piece is known as the agent. It's

interacting with the environment by

taking actions which updates the state

um of of the environment. So that's

that's why you see this word state here.

This gets updated constantly every time

you take an action. Um ultimately what

reinforcement learning is trying to do

is learn the best action like what would

be the best action to take. Um

and the best action is is the one that

leads to the most long-term reward.

That's the best action. Um, so you have

to uh you have to learn what you know

what leads to a good reward by kind of

experiencing this over and over and over

through trial and error. So there's a

lot of um kind of simulation or letting

the robot try something a lot um in

order to kind of learn what's rewarding

and what's not. Think about it again

like I think a good example is like with

children, right? you kind of have to let

them try things until they learn on

their own what's what can they do and

what can they not do

what's the best actions right

so reinforcement learning has made its

way into other places so I I said like a

good example is self-driving cars or ro

robotics a lot of reinforce

reinforcement learning is used there one

place it's found its way into recently

is recommendation systems have kind of

merged with reinforcement learning

learning. Um, and this is because you

you can imagine there's kind of a

built-in reward for you clicking on a

video and kind of watching it.

Um, so that kind of reinforces that

recommendation and then uh that's where

um you can then kind of recommend a

similar thing and see if that's

rewarding and generates a click or

generates some view time or watch time

or whatever. Um so reinforcement

learning has found its way into a lot of

areas. Um recommendations being one of

them because it's just natural for the

idea of like what um should I recommend

next to generate the most reward. In

this case, the reward is kind of

correlated to did they click on it or

not or did they how long did they watch

watch for longer it's more rewarding

um those kind of things but uh place

places where reinforcement learning have

been used I said self-driving cars um

games so uh one of the most famous

examples if you want to look it up is

the um Alph Go this was in 2016 um the

Alph Go uh algorithm was a reinforcement

learning bot that beat um some of the

world's best Go players, which if you're

not familiar, Go is a um board game

that is a little bit more uh complex

than chess. It has more more uh it's a

larger board um more pieces to it. Um

but they there was a reinforcement

learning powered bot that actually um

learned how to play the game so

effective it could beat um world kind of

masters at the games was pretty amazing.

Um that's the alpha go and that was by

deep mind Google and deep mind in 2016.

That was pretty that was only 10 years

ago not that long.

Um so certain uh we said recommendation

uh even autocorrect um learning to

predict like what is the best correction

uh to generate a reward which would be

like you accept that correction or you

reject it would be a penalty. Um so

reinforced learning has been adapted to

these kind of problems very

successfully. Let's take a look at the

packages that we will use throughout. So

um of course we will rely on these three

which we've already relied on to do a

lot of things like numpy to do numerical

manipulations and calculations.

Uh mapplot lib to do any plotting and

not only map lib but maybe seabour as

well both of those to do plotting. Um,

pandas is a big one because

that's where all of our data is going to

be manipulated and prepped before it

goes into modeling.

So all of that stuff we learn from

pandis is definitely going to be applied

here in this course uh as we actually

build models. Um so of course like these

old ones that we've been working with

quite a bit um still going to be useful

here in the modeling stage. Um mainly

for different reasons though mostly to

get our data prepared to do some type of

modeling or maybe to visualize it before

we do modeling to get a sense of what it

looks like those kind of things.

Um,

sci-fi is sometimes useful for certain

uh um processing like in unsupervised

learning. We'll actually use scyp a

little bit to do dimensionality

reduction or help us do that. Um so

scypi will be used here and there and

we've seen it before with hypothesis

testing. We use scypi like the t test

and z test came from there. Um, some of

the unsupervised learning stuff will

come out of there, but the package we

will use by far the most in this course

is going to be Scikitlearn,

which is here. Um, and we've already

seen a little bit about scikitlearn in

terms of its pre-processing capability.

So, we use the uh minmax scaler and the

standard scaler from there from the

pre-processing module in scikitlearn.

but it has um many different models

built into it that we can use to help uh

do our training and predictions. Um so

it's a incredibly useful machine

learning library. It is the industry

standard machine learning library. Um if

you're going to do anything in machine

learning, it would be expected that you

know how to use scikitlearn.

Now what's really lucky about that is

that scikitlearn is a really easy

package to get used to. Nearly

everything we do in scikitlearn will

mostly follow the same pattern and so um

the code will be extremely simple. They

did a great job with that package of

making things really user friendly,

really simple. Um it's a really

fantastic package and we're going to get

a lot of practice with it uh as we go

along. Every model we build will

essentially be from scikitlearn

and not only like the models but um

doing the training doing the predictions

and then doing the evaluation will all

come from different uh scikitlearn u

modules. So that'll be really nice and

we'll get um good exposure to that

package throughout the course. So if

anything will come away from this course

as um psychit learn uh uh experts

that'll be very nice. So this is this

will be the new one for us psychitlearn

but we'll get a lot of practice with it.

Okay.

All right. So just to recap that lesson

before we move on to lesson three. Um we

talked about machine learning as

learning from data. um which is included

underneath the AI umbrella. But deep

learning is also included under machine

learning because it's still learning

from data but it's learning using neural

networks.

Um we talked about the four different

types of machine learning. We had

supervised, unsupervised,

semi-supervised and reinforcement. So

those are the the different types of

machine learning that are out there. Um

and then we talked about some of the pi

python packages uh that we will use the

main one being scikitlearn and of course

we'll use our older like pandas to

manipulate our data and get it uh pass

it into our model training etc.

But scikitlearn will be uh our go-to for

anything machine learning.

All right. So, some questions for you

guys, some checks.

So, let me know in the chat. What do you

guys think? Uh, which of the following

best describes machine learning?

Which choice do you think makes the best

is the best for this?

Very good. Very good. I see I see a lot

of choices for A and A would be the

correct choice. So machine learning is

definitely um a a subset of AI. that's

underneath that AI umbrella, but of

course we're learning from experience

and of course that experience is

recorded in the data um without being

explicitly programmed. Uh so it's the

exact opposite of BNC. We're definitely

not learning from rules and it's

definitely not just used for image and

speech recognition. It can be used for

many other things beyond those. So yeah,

A is the best choice there.

What do we say here?

Okay. What do you guys think about this?

Which example illustrates the use of

machine learning to enhance customer

experience in an ecommerce company?

In other words, what would be some what

would be some uh typical use cases of

machine learning?

Good. So I think uh C is going to be the

best answer here. Definitely C. So it's

using machine learning to do uh fraud

transactions. So so that would be a

prediction probably a supervised

learning right if if this is fraud or

not fraud. Um and then maybe some

customer behavior uh that might be

unsupervised. So maybe grouping together

customers uh clustering them based on

their data like their shopping behavior

and characteristics. Um that that might

be unsupervised but either way it's

machine learning.

Okay.

Okay. Final one. What distinguishes deep

learning from machine learning and

artificial intelligence? So what's

unique about deep learning?

Oh, very good. Yep. So, deep learning

uses neural networks as so you guys are

right on top of that. Neural deep

learning uses neural nets. That's what

makes it unique. So, machine learning

would be part A. Machine learning is

focused on learning from data.

underneath of that is learning from data

using neural networks which is what uh

deep learning is.

Very good.

All right, let's go to lesson three.

And lesson three has two notebooks.

We're going to be starting with 3.1.

So, you'll want to open up that

notebook. I'm going to go over to it

now. Give you a moment to open that up.

So, we're going to open the 3.1

notebook. Um, there's two of them. We'll

see how far if we can get into the

second one today. Probably will.

Um, but we're going to do the uh we're

going to start with 3.1 notebook. Do you

guys have this notebook? Should be in

your materials for for this course.

Let me give you a moment to open that

one.

Do you guys have it?

All right. So, we're going to start by

talking about uh supervised learning

um in our machine learning journey. So

remember, we're going to talk about uh

supervised and unsupervised after we do

supervised. Um and there's going to be a

lot to cover with supervised mainly

because um there are uh two different

types of problems we can tackle uh which

will be uh we'll talk about in a moment

predicting different kinds of values. Um

but let's talk about the kind of what

we're hoping to learn here which is um

talk about the different kinds of

problems that we'll study which are

these these categories of supervised

learning. Um those two categories are

going to be called classification and

regression. We'll talk about those and

their differences and then talk about

some applications and some uh example

algorithms

and that's just within this notebook. Um

3.2 two we'll get into uh regression in

particular

um which will be uh very very

interesting. Okay. So that'll be our

first models that we'll build will be

over there in 3.2.

Okay. So if you guys remember um

supervised learning is where we learn

from labeled data. So we have input and

outputs in our in our data set. Um and

you so you train a model on this data

that includes input features and

corresponding outputs that are that are

the labels. Right? So um the goal is to

learn a relationship between the input

and the output. Of course that's what

any model is trying to do. Um, and what

this allows us to do is then take that

model and use it to make predictions on

never-beforeseen

uh data. Right? So then we have a

predictive model out of that that we can

use um going forward on new examples.

Um so

remember we will have in our data a

bunch of features which are columns and

then generally one of those columns will

be the label that we're trying to

predict.

And our model is going to try to learn

some type of relationship between those

inputs and the output label. So the

output label could be like fraud not

fraud, cancer not cancer, uh a price, a

temperature, those kind of things.

So let's talk about that. inside of um

supervised learning there are two

different types of learning that we can

do and they're really based on the label

or sometimes that label is known as the

target that we're trying to predict. Um

and depending on that type we get these

two different categories of learning or

two different types of learning. One is

known as regression. So that's generally

when we are predicting something that is

continuous or something that is a

numerical.

So numerical

numerical value. So think of price,

think of temperature, think of revenue.

We're trying to predict something like

that. Um versus something that is

categorical. So that the predicting

something categorical would be like

fraud, not fraud, spam, not spam. um

those are discrete categories and the

problem of predicting categories is is

known as classification because we're

trying to classify examples as belonging

to one category or another.

So we have these two main types of

supervised learning problems. we have

regression and we have classification

and they're going to be handled slightly

differently

um for many reasons that we're going to

uncover. Um one of the primary reasons

is that of course we're predicting

something that's continuous in the

regression case versus something

discrete. So the models have to be

slightly different to account for that.

Um but then a step beyond that is the

evaluation has to be different too. Um I

kind of alluded to this last week, but

when you're predicting a regression,

it's very very difficult to to get the

exact numerical answer. So um generally

we don't care about that. Um generally

we don't care about getting exactly uh

we don't care about getting it exactly

right.

um we just care about getting it um

we're just we care about getting it

nearby, getting it close enough. Um

whereas classification, we do care about

getting exactly right because it's a

discrete category. So we're going to be

able to evaluate that a little bit

differently to say did we get the answer

right or wrong. Regression is going to

be did we get close? Um because it's we

assume it's going to be nearly

impossible to predict a a continuous

number. Um, that's very hard to do.

Okay.

So, any questions on

uh that?

Any questions on those two differences?

Let me give you some examples. Maybe

it'll it'll help too.

So, again, the classification is going

to be predicting uh something that's

categorical. regression is going to be

predicting something that is continuous.

So think about trying to predict the

price of a house based on those other

features we talked about before like

square footage, bedrooms, bathrooms, all

those things we predict the price. That

would be a regression problem because

the price is a continuous value.

Let's take a look at an example here.

Um, imagine we were trying to uh predict

the temperature tomorrow. That's going

to be a regression problem, a a

supervised learning kind of regression

problem because we're trying to predict

a numerical temperature.

Okay? And versus a category like a

discrete category would be this would be

a classification. So this is a

regression on the left. This is a

classification

on the right. Classification

um because we are um predicting one of

two categories. Is it just hot or cold?

Now, we're not saying exactly where that

threshold is on what's hot or cold. That

would be a decision on on what we want

to what our discrete categories actually

mean.

But, um we only have two choices, hot or

cold.

versus predicting the entire temperature

which would be um a numerical prediction

of some exact number. Right? So that'd

be a regression and then on the right

would be a classification. Um now again

why is this so different? You can see

the types of predictions we're making

are completely different. One's a

number, one's a category. But again with

evaluation it's like if the if the true

answer in our labels was 84

and we predicted 83 that's a pretty good

result. That's still pretty close.

That's pretty close to this. So from an

evaluation perspective that's pretty

good. Um whereas like if I predicted

cold and it's actually hot that's that's

a wrong answer. So they're evaluated

slightly different.

Um, and that's something we're going to

see as we talk about evaluation of our

models once we build them is depending

on if it's classification regression,

there's going to be different ways of

evaluating them.

You can kind of see why it's very

difficult to say, okay, we got exactly

84 when it could be any number. Our

model is going to be predicting a

number. That's really hard to pin down

an exact floatingoint number. So, the

best we can do is kind of say, how close

did I get? Like, this would be a worse

answer. If I got something all the way

down here, that's a really long distance

to here. That's bad. That's a bad

prediction. But if I get something

really close, that's better, right?

That's a decent prediction because it's

pretty close,

right?

Of course, being perfect would be

getting exactly right, but that would be

nearly impossible to do.

Okay.

All right. Any questions on this?

Does it make sense on regression versus

classification? We're going to use those

words quite a bit as we go along. So

regression predicting that continuous

value classification predicting a

category

and they're going to be um different

models that do that

different models being used for

regression versus different models being

used for classification.

All right, let's talk about supervised

learning. uh applications here. So just

to name a few, we have HR operations.

Imagine your recruiter tasked with

finding the best candidates. Um so

supervised learning can help by um

rejecting or accepting candidates. Now

this is something that happens quite a

bit even today. Um and that it's kind of

like uh how recommendations happen like

this this resume should be um

recommended this should not um from a

whole pool of applications. Um so

there's those kind of use cases of of um

predicting a category that would be like

a classification. Should we should we

accept or reject the the candidate?

um finance. You see this all the time

with things like risk and loan

approvals.

Um you can uh predict the the the

category of like if the if the loan if

we should accept or reject the loan

application. Um you know that would be a

classification.

Um what's interesting about

classifications by the way so it says

here like we can predict the likelihood

of a of a loan being repaid.

um is a lot of classifications um we we

say that they predict a category but

under the hood they can actually predict

a probability and we turn that

probability into a category. So um you

know like we could say what's we could

say the likelihood of her loan being

repaid is very low. Let's say it's less

than 50% probability. Um then we could

label this as reject,

right? Right? We could label that as a

rejection. Um if it's greater than 50%.

Then we could label this as accept. So

we can set a threshold there

and say okay truly we're predicting a

prob like our model spits out a

probability but we turn that into a

category by saying should we accept if

it's less than 50% we should reject if

it's greater than we should accept.

Okay. So that's something we will see

with some of our classification models

is that they actually produce a

probability and we turn that probability

into a category label

um by by doing something simple like

this putting a threshold on it um for

the for the category.

So finances is used all over the place.

Not only just loans like fraud, we

talked about fraud, not fraud. That

would be a classification.

Um predicting sales revenue, that would

be a regression, right? What is the

revenue going to be in the next two

quarters? That's going to be a

regression problem.

Uh emails like spam, not spam, that's

going to be a classification.

um that's going to operate on the that's

going to take the text input and predict

if this email is a spam or a not spam.

That's going to be a uh supervised

learning problem, but it's going to be a

classification problem,

right? Uh manufacturing supervised

learning is used to inspect and uh

quality and classify products in

different grades. For example, a factory

might use a model to check for defects.

So this is actually something that

happens is you look at images of

products as they go through the assembly

line and you can take a look at those

images and predict if it's a high

quality, low quality, medium quality. Um

so they can be this is a classification,

right? They're going into different

categories of quality. Um so it's much

much like a manual kind of intervention

by some uh QA or quality control uh

specialist.

Okay. But that's a classification.

So in the maritime industry, supervised

learning can be used to predict current.

So current level

um and that can be used to forecast uh

supply and demand. Um so those would be

like regression models that are used to

predict um kind of like temperature but

in this case like title levels.

We talked about fraud already, so that's

there. Um, that would be a

classification.

Okay,

any questions on these uh examples?

Of course, there's many more. Um

recommendation is kind of like a

supervised learning problem uh where you

are

taking examples of things that people

have viewed in the past or or reviewed

in the past and using that to predict

what they would want to watch in the

future. Um so recommendation is

supervised learning. Um and it's like a

classification, you know, trying to

predict um uh certain number of

categories of of uh shows or movies that

you would want to watch. Um

and that's something that we will study

in the future. Recommend we'll we'll

have a whole lesson dedicated to

recommendation as well.

All right.

So when it comes down to the uh actual

models themselves, so there's going to

be lots of different models that we are

going to cover. Um and they are um going

to be different in their purpose and

kind of their uh what kinds of problems

they're used for. Um and uh their their

how they actually train is going to be

different. Um, but at a high level,

they're all trying to do the same thing,

which is learn some sort of relationship

between the input data and the and the

label, right? That's really what they're

trying to do because they're all

supervised. They're they have those

labels, trying to build some

relationship there. Um, they just do it

differently.

And what we're going to study is the

pros and cons of a lot of these models,

like when would I use one of them, when

would I use another. Um, so we'll try to

talk about that as we go along. Um, but

they're all trying to learn some

relationship between the input features

and the output, right? So you have to

keep that in mind. They're trying to

model that relationship. They just do it

in different ways. Okay? So as we go

along and learn about new models, um, we

will learn the details. will learn the

ins and outs um and those pros and cons,

but they're no matter what, they're all

trying to uh learn that relationship,

right? And be able to make predictions

on new data.

Okay,

so here's a list of models that we will

cover and work on throughout the uh the

sessions that we have. um we're not

going to do them all in one one sitting,

but um the first one that we're going to

start with and that we'll cover today is

going to be linear regression.

So we will cover linear regression and

then we'll cover the rest of these guys

mostly in the context of uh

classification.

So, um, what's interesting is some of

these guys can actually be used for both

regression and classification as long as

you make, um, certain adjustments to

them. They have variations that can be

used to do classification and regression

is very interesting. Um but we're going

to start with linear regression today

and then work our way through the rest

of these models when we do um we're

going to do a separate lesson four on

classification. So these all these guys

will come from lesson four.

Um and then uh we will do this guy in

lesson three in the 3.2 notebook. We'll

do all about linear regression.

Yeah, I so logistic regression is a

classification um which is kind of

strange that its name is regression but

it's doing a classification but the the

reason is that the logistic regression

um computes a probability. So it does a

regression to predict a number but that

number is actually a probability. So it

it produces a result that's between it

produces a probability that's between um

obviously uh zero and one.

So it uh and then we take that

probability and we turn it into a

category

like a spam not spam fraud not fraud.

Um but so so logistic regression is kind

of special. It's sort of like a

regression but it's predicting a very

specific type of value which is a

probability. So for for that reason it's

a classification uh algorithm primarily.

So we'll study that one in lesson four.

Uh but yeah, that's that's why it's

under that kind of umbrella of

classification is because it's it's

producing a probability as its main

output which we can then turn into a

category as long as we interpret that

probability as um in the right way uh

like the probability of spam,

probability of not spam.

Okay.

Okay. So, let me focus on um

let me focus on linear regression. I'm

not going to go through all of these

other use cases because we haven't

learned these models yet. Um so, I don't

think they're good. Uh I don't think

it's good to read about them yet until

we've covered them. So, once we cover

them in lesson four, I'll come back and

describe these examples to you guys and

we'll see why it makes sense. But I

think for a linear regression um which

is what we'll cover next, let me talk

about that example. So a prototypical

example would be like predicting the

house prices that we've seen in that

house price data set.

So um if we wanted to uh if we wanted to

predict um if we wanted to estimate the

market value of a house so the price

um we could do that by using the

features such as number of bedrooms,

square footage, location, age of the

property. Um and you know then when a

new when a new house comes on the market

we could estimate what the price should

be based on those features. So linear

regression is a good one to predict the

price like a housing price. Um and we'll

actually practice that in the next uh

notebook.

So we'll we'll uh and then all these

other now there's descriptions of these

other models but again we haven't

covered these guys yet. So I don't want

to really go through those until we get

to those models. So we get to those I'll

come back and mention the example.

Uh can K andN be used for clustering?

No. So um the clustering model is going

to be different. It's going to be uh K

means

K means that's the primary clustering

model. Not K nearest neighbors. K

nearest neighbors is used for uh it can

be used for regression. It can be used

for classification.

So we'll we'll talk about K andN which

is the K nearest neighbors in lesson

four.

It sounds really similar. Yeah, it

sounds really similar but K means is a

clustering algorithm that's that's

slightly different

different uh there's no labels used at

all. This K nearest neighbors is a is a

supervised learning algorithm. It uses

uh labels.

Good. Any any other questions so far?

Okay.

So that being said, let's move on to the

3.2 notebook.

Let's move on to that which will be our

um first discussion around uh

regression. So going into supervised

learning and regression. Give you guys a

moment to pull up this notebook.

But yeah, you want to pull up the 3.2.

We'll do this one next. So we'll focus

in. And so our plan is to do regression

first and then we'll talk about

classification in lesson four

which we will cover all those other

models which you you could use for

classification uh on that list but then

we're going to talk about linear

regression uh first.

All right. So, we have a a big agenda.

This is a big notebook um to go through

a lot of material here surrounding

regression. So, we're we're going to

start with linear regression and see um

how we actually perform it, what that

model is doing. Um which we've kind of

seen the idea of it a little bit

already, so it should be somewhat

familiar. Um and then we'll talk about

how to adapt that linear regression idea

to um nonlinear what's called nonlinear

regression which is going to be using

like polomial

uh features. We'll talk about how to do

that. Um and then a big big big topic

for us is going to be evaluating the

model. So it'll be it'll be quite easy

to actually build it. building the model

will be really easy but evaluating and

interpreting that will be uh a lot of

interesting work there um because we

want to know what the performance of

that model is once we have it built

right we want to know how good of a

model is it is it worth using or do we

need to retrain it or get new data or

change the model up to talk about that

um how do you determine what to do based

on that performance

um and then we'll talk about here um a

couple things. We may not get to this

today, but regularization

which is used to boost the performance

uh in certain situations um whenever the

model is kind of uh performing um poorly

against test data even though it

performs pretty well on training data.

In that scenario, you can use offshoots

of linear regression that do some uh

what's called regularization. We'll talk

about that.

Um and then we'll talk about

hyperparameter tuning uh generally as a

strategy which is something you

generally do want to do when you're

training machine learning models. Um so

again these two we may not get to today

but um quite a quite a lot to get to be

prior to that mainly centered around

evaluation and building linear

regression.

Okay. So pretty cool. we'll get to our

first kind of model here. This linear

regression

to start with.

Okay,

so let's start with uh linear regression

here. Um, and really what linear

regression is attempting to do and I

want to show you this in this picture is

draw this line sometimes what is known

as the line of best fit. So this is our

model that kind of goes through the data

and it's generally a good predictor

um because if you give me um features uh

if you give me new features and let's

say they are let's say you give me a

feature that's right here.

So you say, okay, I have a feature

that's this value on the x- axis. Then I

know all I have to do is plug that into

my line equation, and I will generate a

a value that's like right here.

Okay, that's pretty that's on that line

at that input. And that's going to be my

prediction for what the output variable

should be. It's just going to be

something on that line. And what you can

see is this line is a decent estimate

for this data because it slices through

this pretty evenly. So it's a good guess

as to what the output should be given

any one of these inputs. It's a it's a

good estimator this line. And so our

goal building a linear regression is to

kind of build the equation of this line.

So we want this equation.

Equation of this line

is going to be our model.

Yes, it's going to look just like that.

MX plus B or yeah, MX plus C. It's going

to look exactly like that. uh except

that it's going to be more than just MX

because we have um generally more than

one feature. So you think of X as a

feature um it will be more than just MX.

It will generally be like uh it'll

generally look like this

and then plus maybe some bias here plus

uh an intercept. Yeah, it'll generally

look like that. So, yeah, you're exactly

right. MX plusb is the right idea.

Exactly right.

It'll generally look like that.

Nonlinear. It can be adapted to

nonlinear. Yeah. If we transform, we're

going to talk about that. If we

transform all of our features in a

nonlinear way, um we can apply linear

regression to it. Yes. And and that

would be a nonlinear regression. So yes,

we can do nonlinear things too.

We'll talk about that.

Okay. So linear regression again is the

art or science I should say not really

art but it is an exact science of

finding the equation of this line that

fits through this data. Um now why one

thing you should be thinking about is

why is this line a good predictor and

the argument is that if you take a look

at this distance from these blue points

so let's say these blue points are our

actual data points this line is going to

be found such that it minimizes this

distance

from the points to actually I should

draw it this way from the points to the

line.

So, we want this distance to be um

actually I should draw it that way, this

way. We want this distance to be kind of

at a minimum. So, it would be bad to

draw a line all the way out here because

then that's a lot of distance, right?

So, and that would be a lot of error um

contributed from not being able to

predict those points in our data set

very well. Um which is our training

data. That's why we have labels, right?

that that guide us in building this

line. Um so our goal is to build that

line especially so that this error or

this distance can be as minimum as

possible. Right? Which are all these

distances from these points to the line.

We want those to be as minimum as

possible. So our goal is to find this

equation.

So we're going to build a model that's

going to find this equation.

of the line

um such that our error

is minimal.

And what is the error? The error is the

distance

of our data points

to

to the line that we build. So

essentially what we'll do in order to

train this will be to adjust the

parameters or the or in that like I

think it's really good you brought up

the MX plus C. Basically the M and the C

will adjust. So we adjust those

accordingly to make this distance as

small as possible.

Okay to minimize that distance as much

as possible.

Okay.

So, um where is regression used? We've

already seen some examples. Here's some

more uh advertising like predicting

sales, predicting um oil and uh oil

production and demand. Those are like

forecast those are regression problems.

Um retail like demand forecasting for

inventory. Um healthc care predicting um

uh the levels of certain um uh blood

markers or you know something like that.

Um real estate predicting prices based

on those uh talked about like square

footage, bedrooms, bathrooms, those

things. So regression is used again

whenever we want to predict a number a

numerical output um that's a regression

problem.

So this kind of regression we're talking

about here is generally

um known as uh a when that equation is

linear that is known as a linear

regression. And so go back to that

picture when we have a when that

equation of the line that we find is a

linear equation meaning that it is

exactly the form I've been telling you.

So it's it's something like um weight

time feature

plus weight time feature

plus weight time feature

and then maybe some intercept um term

like some some bias term there. Um this

is a linear equation because all of the

features are to the single power. So

it's a linear power and this is a linear

combination of features with with those

different weights. So this is a linear

model

because it is uh it's what in math we

would call this a linear equation right

everything is to the first power. It

resembles mx plus b. It is a linear

equation or a linear model. Um so when

we talk about linear regression that is

a regression model so we're predicting

some continuous target that assumes we

are model our model is formed from this

kind of equation a linear equation.

So this is going to be our our model for

a linear

uh regression.

Okay.

And so when you when you train a linear

regression, your goal is to learn these

weights so that you can plug in um you

can plug in any one of your uh input

features and you um can generate a

prediction. You can which is going to be

something on that line, right? It's

going to be a value that's sitting here

on this line.

We put in all of our features and we end

up there somewhere on that line.

This output.

Okay.

Okay. Let me pause there. Any questions

on the linear model here or why it's

called linear regression?

Okay. And by the way in these notes um

this bullet point here where it says it

uses the least squares criterion to

estimate the coefficients that is

exactly what I said earlier with the

distance. So the distance is based on

the square

of this this quantity like how far away

you are from the line is based on this

square distance here and here and here

and here. So what we're trying to do is

find the least distance or least squares

which is that minimum distance. So

that's how we find all of these weights

is from minimize. We basically tune them

enough using our labels. So here's our

label which is the y. We basically plug

in our data and tune those enough to

minimize the error. It's it's a it's an

optimization problem,

right? We we're trying to find the

minimum of this quantity which is that

best fit line.

Okay.

So we have linear regression

um and we can do a simple linear

regression that only has one feature. So

if it only has one feature that's

exactly the so if there's only one input

feature sometimes that is known as um

simple regression or simple linear

regression and there's basically there's

only one feature. So one independent

variable is the feature.

There's only one feature. And so this

equation resembles the

exact equation that you guys just put in

there, which is um mx plus b,

right? It resembles exactly that. Um

we're just using different symbols for

those like beta beta 0 and beta 1. But

um basically exactly that simple line is

only one feature. So, and that's because

that line is going to um that line is

going to be generated uh according to

that equation. So, here's kind of what

it looks like.

This is the best fit line through all of

these blue dots. This is something we're

going to be able to build. We're going

to be able to build that equation um

pretty easily in scikitlearn.

So, we'll be able to find that um and it

won't be too hard. So this line will be

um y = beta 0 plus beta 1. So some

weight beta 1 times the only feature we

have x1.

Okay. So in this case um we would be

predicting sales. So sales would be the

value basically the label that we're

trying to predict and the feature that

we're putting in is uh I think it's the

number of TV expenses. Yep. TV expenses

which is on the x- axis. So there's one

feature which is um TV expense.

So um on this graph this would be this

would be our model.

Okay that would be our model. We only

have one feature and we have um these

two weights. We have an intercept B 0

and or beta 0 and then a one weight

which gets applied to that one feature

beta 1. And so our model would have

certain value for beta 0 and a certain

value for beta 1. That's what get that's

these guys get learned

learned during

model

training.

Okay. So those are what get learned

during our model training and they get

learned by a a a least what's called a

lease squares algorithm that is trying

to minimize that distance. It tries to

tweak beta 0 beta 1 to minimize this

distance of this line

um this line

to all of these points

trying to minimize this.

So imagine taking a line and kind of

moving it around and turning its its

slope, its angle um to try to find that

best fit,

which reduces that error the most.

Right? That's kind of what we're doing.

Uh can I explain? Yeah. So uh sales is

in dollars and and TV expense

um

uh

TV actually I think it's the other way

around. I think the sales is actually a

quantity. So I this is number of sales

that we have and TV expense is um I

think I think it's in dollars. So how

much money how much expense um did we

put into the into the product and then

this is how many sales did we have of

that product.

So I think it's the other way around

but what this what this graph is showing

is the blue points are our actual data

points. Okay. So so we have a collection

like we have a data frame that has so

imagine we had a data frame that has the

uh true values.

So it has the um TV expenses.

Um it has points that are like one. So

it has points that are like 120 and then

the sale sales could be like 700

700 units, let's say. And then it has um

so this is just our data set, right?

This would be like in a data frame that

we have. And then we had ones that were

um 50 and then this could be um this

could be 400 let's say and on and on and

on right so this is our data and this

data is plotted in the blue so these are

these blue points here

right so these are the blue points here

and the red points are is our model so

we built a linear regression model um

where we are putting in some values

we're putting in some fake x values

here and generating some predictions

which is this line,

this linear uh regression line, right?

And that line is derived from this data,

right? It gets learned from this

supervised uh examples.

Does that make sense?

That line is derived from the data. it's

actually um learned from like the line

of best fit is learned from that data

and the actual data is in the blue.

So you can see we're trying to build

this such that this distance is kind of

a minimum

so it's an optimal fit

to balance out these distances.

So it's just plotting. So it's just

building that relationship between the

input and output. Like when the when the

expenses are higher, um we seem to have

more sales.

Okay.

Uh what's perpendicular like the

distance? This should be this should be

perpendicular because it's a distance

here.

Is that what you mean? Like the distance

from the real points to the line? Yeah,

that should be perpendicular

because it's it's a it's a distance

formula.

Okay.

All right. So more generally now do do

we usually have one feature? No. So

generally we expand this to the more

general case where we have more than one

feature like what we see in the housing

data right where we could predict a

price but we have many different inputs

like bedrooms, bathrooms, square footage

etc.

So more broadly

instead of simple linear regression we

have what's known as multiple linear

linear regression which means we have

multiple variables or multiple features.

Um so this is exactly the equation I've

been talking about. Um so we just extend

that that one into many features. So

which is this case and then a intercept

term which is uh um there as sometimes

known as the bias. Um

but this is the intercept term to kind

of orient the line to start out in the

right place. Um and uh but this is the

um this is the equation that we would be

building the model. This is our model

essentially, right? This is the equation

we would be learning.

Intercept is like a constant. Yeah. So

if if all of the features were zero, um

this is what our our data would be. This

is what our result would be. If

basically if this was zero, this was

zero, this was zero, it would reduce to

this as the prediction. Yeah. It's like

a constant. Yes.

So in in geometry, the intercept's

actually really important because it it

orients where your line should start. So

it orients like so so these values are

kind of like the slope. They orient the

tilt of it. Like should it be tilted

like this or should it be more sloped?

But the intercept orients where it

should start like vertically like should

it start all the way up here? Should it

start more down here?

Um, that's what the intercept kind of

tells us.

Okay, so this is the situation. This is

going to be our linear regression model

that we will be building most of the

time because we will have again these

are all going to be features.

So this is some feature the X this is

some feature this is some feature

X1 etc. These are all features and what

gets learned during the training are

these coefficients. So all of these

coefficients including the beta 0ero um

will get learned. So these will get

learned

um from our data right they get learned

they will be trained from our data um in

order and and how do they get trained

it's from reducing that distance we try

to get that line of best fit by tweaking

those betas enough to uh until we reach

a minimum distance but there's there's

an algorithm behind that um that that

scikitlearn will run for us to find that

best fit Um, so we don't need to do that

manually, but that's that's the process

is basically tweaking those weights to

end up with that line of best fit. So in

higher dimensions, instead of a line,

you get more of what's called a plane

here. Um, which kind of looks like this.

So the best fit is actually this plane

where all um, it kind of dissects all

these points just like that um, in

higher dimensions. So this is uh instead

of a line you get this in in three

dimensions you get this plane like this

but it's still it's like a line of best

it's just a more general line of best

fit. It's still the same idea. Um we're

still trying to um come up with the best

coefficients to minimize that distance

from our from our points to the line.

Although in higher dimensions it's no

longer a line. It's more like a plane

like this. So you're trying to minimize

this distance from here down to the

plane

here up to the plane

in higher dimensions. So I want you to

keep in mind what we're trying to do

before we go into the code because the

code's going to make it seem really

really simple and that's because

scikitlearn is great and that's what it

does.

But we should realize that there's

something really complex going on which

is again finding the best value of these

weights

that minimizes the distance of this line

to the data points that we have. So

there's an algorithm there that will

keep trying to make adjustments to this

based on those distances. So it's going

to use those distances as a guide to

kind of tweak them to find the one that

results in the lowest amount of

distance. So we keep making tweaks, keep

making tweaks, keep making tweaks and

eventually we try to find we converge to

the set of weights that gives us that

best fitting line. Um and and there's an

algorithm there that occurs. Now luckily

that gets abstracted for us a bit behind

um scikitlearn

um finding that best fit. So there'll be

a function that we use in scikitlearn

when we build the model that will go

ahead and find the best weights for us

and that's then we now have our optimal

model right that then we can just plug

in different values of these features

and generate a prediction which is going

to be this uh result right so so that's

what we're ultimately trying to do is uh

train the model which will uh find all

those optimal weights and then uh we can

predict with it which would be plugging

in different feature values to to

generate a prediction.

Okay,

so let's see how that happens. It's

actually going to be super easy um with

scikitlearn.

So uh in this scenario we have um we're

going to import our pandas because we're

going to load our data from that. Um, so

of course we need some data to work

with. So we're going to load this uh

CSV.

Um, I

uh so I was not actually able to find

this CSV for this example, but I mean

that's okay because we'll do some we'll

do other examples where we'll work with

the data. If you happen to have it, um,

great. I didn't see it in in my files.

So just have to take the word for it

that these are the this is that TV and

sales columns here um from this data

set.

Okay. Um as an example. So um just to

see how it's fit um what we're going to

do and this is going to be a very

standard process for us for building a

model. These steps are going to be very

very standard for us which is going to

be first of all splitting the features

away from the label. That's the first

step that we always will take. So if you

take a look at this code, it's taking

all rows but only the first column.

Okay, so it's extracting all the

features from the data frame um which

happen to be which is just the first the

first column uh which is the TV uh

column right just that column there and

our target variable which is our label.

So our target variable aka the label um

is the second column, right? It's that

that sales column.

Um and so our first step here, let me

call that out here. First step is to

always split apart

features from labels.

Okay, so we put all those features into

a data frame called X and we have all of

our labels into technically a series but

uh sort of like a data frame, right? Um

called Y, which is just the um which is

just the uh uh labels. So that's just

the TV values. Um now you're going to

see why we do that. It's because we need

um our our features and labels split

apart to put them into the model

building function. It expects our

independent variables or our features to

be separated from our answers or our

labels that guide the model building.

That's the first thing you got to do is

separate those.

Okay, so this code will separate those

out into a capital X and a lowercase Y.

And that's actually pretty industry

standard notation. Whenever you split

apart all your features, usually you put

them into a data frame called capital X

and then you have a lowercase Y to

represent your labels. That's actually

pretty standard.

So it's pretty standard that um X

represents

features

and

Y represents labels

label column

whatever our label column is in this

case it is the sales because we're going

to be predicting sales

using the TV column the TV quant expense

quantity.

Yeah. So what it so the assignment is

that we are um the assignment is that we

are

uh we are um splitting apart our data.

So that when we first read in the data

um it is a data frame right that has two

columns TV and sales.

Oh perfect thank you Tim. I will I will

go ahead and so if we look at this data

it only has those two columns right it

only has those two columns. Okay. So

what we're doing with this is we are

splitting apart

our our independent variable our

features. So this this x will contain

our features

and y will contain

our label.

Does that make sense? We're splitting

this data apart. So, we're only grabbing

that first column here to be our

features. And then we're we're grabbing

the second column, which is the sales,

because we're going to predict the

sales. This is our label. We're going to

we're going to build a model to predict

the sales given the TV input, TV expense

input. So, the first thing we have to do

is split apart the features and the

label.

Okay, that's the first step we usually

will take. And the reason we have to do

that um just to reiterate, the reason we

have to do that is because our model

will expect our our data features to be

separate from the label. We will pass

those in separately.

X is TV. It's the first column

because we're using eyeling.

We are predicting the sales given the TV

expense value.

Yeah. which is why we split it into so

this is the second column right the

index one column

uh you just put in read CSV and pass in

the URL so you could so exactly the code

that was up earlier from Tim

um you just do this

and then data equals ed read CSV URL

So we split our data into X and Y here.

All right. Now, one other step that

we're going to take that's a very very

critical step and you're going to we're

going to see this step over and over and

over and over again. So splitting apart

into X and Y will become we'll do that

over and over and over and over again.

Not only that, but doing this next step,

which is what's called a train test

split. Now, let me show you what the

train test split does. It takes our data

and it's going to split apart our data

that we have, our X and our Y data. It's

going to split it apart into a

percentage that will be used to train

the data

and then a percentage that will be used

to test. Now, why would we want to do

that? It's mainly so we can do

evaluation. So, we build the model over

here and then we test it on data that

has not seen before. So, we reserve a

percentage of the data to be used for

test. Usually this this data is um

somewhere between uh 20 to 30%.

So somewhere between 20 to 30% of the

original data. So that means the

majority of it is used for training. So

the majority of the of that X and Y over

here is going to be between 70 to 80%.

will generally be used for for uh for

training. Okay. So somewhere between 20

to 30 the industry standard is some

anywhere in between there. Um a lot of

people like to use 30%, some people like

to use 20%. Um anything in that range is

acceptable. Um we will I think we

generally will favor like 30%.

um to be used for testing. But um the

the point is we don't we don't want to

mix those together. We want those to be

separated out so that we can have a fair

evaluation, right? We want to train our

data on this train our model on this

data and then see how well it performs

on this data that it has never seen

before.

Right? So in order to have data it's

never seen before, we're going to take

our x and our y and we're going to split

it using this function called train test

split that will do this kind of

splitting for us. Okay, so scikitlearn

has a function called train test split

that will go ahead and we're going to

pass our x and our y and we'll pass in a

percentage like 30% that we want to

split out into a test set and then the

remainder of that the 70% will be used

for training the model.

Okay.

So what we're going to get let me redraw

that. So what we're going to get out of

this for the train test split is we're

going to we're going to have an X and a

Y per

training and test. So we're going to get

now we're going to get an X train

and a Y train.

So we're going to get training features

and training labels. And then we're

going to get test features

to plug into our model and and test

answers or test labels

to do evaluation because what we should

be able to do is build the model over

here and then apply the model on this

data. Meaning we can take these features

and plug it into our model and then see

what answers we get and compare those

answers to this testing data. Right? We

should be able to do that to generate an

evaluation.

Okay. Now you may be wondering why do we

do any of that? What's the purpose of

that?

Evaluating it on this test data gives us

a good sense of will our model

generalize to new examples. Right? If it

performs pretty well on this data,

that's a good signal like when it's

performing pretty well on data it's

never seen before, that's a good

indicator that it's going to perform

pretty well when we use it on brand new

examples

um in the future.

Right. So that's a that's why we do this

evaluation on this data that it has not

seen before. It's going to see this

training data, right? We're going to

train the model on that data. But that

model will never be exposed to this test

data until we do the evaluation

and and generate some metrics to see how

good is this performing

and does it have a good chance of

generalizing to never before seen

examples which is what we want right

because we're going to use this model in

the real world. It's going to be being

used on new examples that it hasn't seen

before. We want it to perform well. So,

this is kind of our test, our

evaluation.

Okay. Any questions on the We're going

to do this in a moment. I'll show you

what it looks like in the code, but any

conceptually any questions on the train

test split idea. It's a very very

important idea that we um basically use

part of the data to train it and then

another part of it to evaluate. It's

very important we do that. By the way,

this has a term um this in machine

learning this is called cross

validation

because we are using one data set to

train the model and then we're cross

over we're crossing that over into

another data set to validate it which is

the uh the the testing that.

So this is called cross validation. Um

there's actually many ways to do cross

validation. That's something we'll

study. This is a very simple way of

doing cross validation. There's more

complex ways. You can take your data and

you can actually divide it into many

sections

and basically train it against most of

these and evaluate it against one at a

time and then rotate. So that's another

way to do cross validation. We're going

to study that. Um but this is the this

is the simplest way to do it here.

Okay.

So let me show you what you get when you

use train test split. So uh we're going

to import from sklearn.

We're uh from the model selection

module. Now we haven't used this before.

This is our first time using it. But

here's our model selection. We're going

to import this train test split function

and we're going to use it on our X and Y

and we're going to set a test size of

30% which is which is.3. So our test

size

is 30%.

Converted to decimal

right converted to.3 so that means we're

reserving 30% for that test set. Um you

can set a random state. Now that's

completely optional. Um the random state

is for reproducibility

because what the train test split is

going to do is it's actually going to

shuffle the data and then split it apart

into the 7030.

So um yes, the seed. Exactly. It's like

a seed. So it's it's saying like when

you do that shuffling every time I run

this notebook I'm going to get the same

result but it's going to be random the

first it's going to be random but I'm

going to be able to reproduce that

randomness with that random state. Yes,

it is like a seed.

Uh it's you can choose any number to be

your your um your random state. It 42

isn't important. You could choose zero.

You could choose one. Um, you could

choose any positive integer. Um, 42 is

kind of like the uh industry standard.

It's it's you'd have to look it up why

it is. Um, apparently 42 is a special

number. Um,

in kind of the history of development of

this stuff, there's nothing really

special about 42. You could choose a

random You could choose a random seed to

be uh zero. That's fine. It it doesn't

really it doesn't really matter.

Um you just want you can choose it to be

uh one, two, three. Um you can choose it

to be 15. You can choose it to be

anything you want it to be. It's really

so that your your shuffling is

consistent. Every time you run this

notebook, you get the same shuffle

result. So I'm always going to get the

same rows in these splits.

Hitch. There it is. I knew it was from

something.

Yeah. So 42 is kind of like a

it's it's just used ubiquitously

uh you know as kind of a um paying

tribute to the Hitchhiker's Guide to the

Galaxy, but it's no it's there's nothing

that special about 42. It doesn't it's

not going to change our result or

anything.

It's just so that this train set split

is going to shuffle our data and split

it apart into 7030.

You just want to set this to something

so that you get a cons every time we run

this notebook, we get a consistent

shuffle.

And so the data in these sets

are uh consistent. That's all.

Okay. But do you guys see how we pass in

our X and our Y and we generate four we

generate four different data uh

quantities here which is we generate

training features, test features,

training labels and test labels because

again we are generating these four

different we're generating data on these

two different sets a training set

and a test set. So we have training

features, training label,

and then test features, test label.

Okay, that's why it's so important to

split apart our data into the X and the

Y. We need those split apart in order

for this part to work.

So by the way, these two steps we will

always do for any model we build. We'll

generally do X and Y and then train test

split in order to generate the data that

we will use for building our model.

Okay. So this this data here is going to

be what we actually use to guide the

training of our model. So it's

definitely supervised, right? Linear

regression

um we we will use that

Okay, so we haven't built the model yet.

We're just getting our data split apart

and ready for the training. We haven't

actually built our model yet, right?

That'll be coming up uh in a moment. But

this is getting our data ready. We

started with our data frame. We split it

apart into uh an x and a y. And we split

that into a train test split. And um you

know then we can uh then we can go ahead

and um pass in to our model training

which we'll do in a moment.

Um you that's a good question. You could

run so what you could do is you could

run

um should we import numpy? Let's see.

We did. Okay. You could run the average

on the um you could check the MP mean on

the X train and see how it compares to

um

see how it compares to X.

So you could you could do that and see

what the average of this feature is um

compared to the average of the original.

They may not be perfect because we are

taking a reduced data set size. So I

don't think there's really any good

there's not like a one-sizefits-all

validation we can do because we're

taking a random shuffle and taking a

percent. We're taking 70% of the data

out. So we're not guaranteed to maintain

the same statistics. We can see if

they're close.

Um but does that make sense? Like we're

not guaranteed to get the same stats

because we're taking a slice of it.

We're taking 70%.

So it's not guaranteed to to to

be the same distribution really.

Delete that.

Uh is it good practice? Yes, it is.

It is. Uh 30% is the industry standard.

Anything between 20 to 30, so 0.2,

0.25.3,

any of those are acceptable. It's really

up to you. Um I mostly see 30%.

Mo I think.3 is is a good good practice

to use for sure.

Um I did explain random state. Uh random

state is so that you get consistent

shuffling. Um you can set this to any

integer that you want it to be. It it

doesn't really matter. Um you can set it

to uh 100, you can set it to 10, you can

set it to 15. Um it just ensures because

what this split will do is it will

shuffle the data first. It'll shuffle

the rows and then um split it apart into

the into the train and test sets. So you

set the random state so that the next

time you run this you get the same

consistent shuffling. That's the only

that's the only thing it it helps you

with because it is randomized but when

you set a random state um it's so that

like if you run it again you'll get the

same shuffling.

You'll get the same the shuffling

matters because it it it uh dictates

what ends up in in these sets.

Okay.

All right. So let's see let's do let's

build the model. Um and let me show you

how easy this is going to be to build

the model. And this is really how it's

going to be for every single scikitlearn

model will basically look the exact same

for training it which is what's going to

make it really really nice. So the first

thing we have to do is import our model.

So from scikitlearn we're going to be

using a linear from the linear model

package or the linear model module I

should say within sklearn we're going to

be importing the linear regression

and we're going to create an instance of

the linear regression here.

Okay, so linear regression and look how

easy this is going to be. Nearly all

nearly all sklearn models use

ffit function to train.

So every one of them, no matter which

one we use, like the decision tree, like

the um logistic regression, any of those

like we use for classification that are

going to be coming up in lesson four,

they're all going to look the same in

terms of it's going to run.fit,

which is um scikitlearn's

uh generic function for training your

model. So this will execute the training

once we run this code. And what that

again the linear regression training is

going to do that least squares distance

procedure or algorithm to try to find

the right weights. It's trying to find

those weights that minimize that squared

distance uh from our line that it's

trying to build to the data.

And what I want you to notice is what we

put into the ffit. See how we put in the

training data where we put in the

training features and we put in the

training labels. Now this is supervised.

So of course we put in the labels,

right? Of course we put in these labels

here and of course we put in our

features here. So we're putting in all

of our examples from our training split

into this ffit which is going to train

the model uh so that we can we can use

it for prediction.

Okay, it's really fast. If I run this,

it's going to be pretty much instant.

Pretty much instantly it gets trained.

And you can see here we now have a

linear regression. you can see in this

little box. Um, and it and this

information says that it has been

fitted. So, it's now ready to be used.

Right? So, we now that's it. We've

trained our model. We try that's how

easy that was. We did ffit. Now, what we

should realize is there's a lot of work

going on behind the scenes of this ffit.

Okay. There's a lot of work being done

there to do the least squares algorithm

and find those weights and and create

that line of best fit. Right? So there

there's a lot of work being going on

there that's going on there behind the

scenes, but scikitlearn is abstracting

it away for us, right? And all we have

to do is fit when we're using this code.

Really easy. Really easy. Fit. And there

we go. We've trained our linear

regression model.

And by the way, if you want to see what

the coefficients are, you can actually

extract them if you do so if you take

your lin regression and you do um

coefficients like this.

COF with a with an underscore. So this

gives us the trained

weights

coefficients

also known as the coefficients right.

Um so if you run this you can see uh

right now we have this coefficient here

um which is the only coefficient we had

on our feature. So we only had one

feature coefficient there.

And we can take a look at our intercept

which is this.

So this gives us the train weights

and so we can look at the intercept we

can look at the the the coefficient. Um

so obviously if we have multiple

features our model has many features

it's going to have more values in that

coefficient but the intercept is just

the single value 7.23

and then the coefficient

is 0.046. So that's the weight that gets

learned.

Is there a size limit? No, not really.

There's no size limit. Um,

no. You can use as much data as you

want.

There's really no size limit other than

what like what you can fit in memory.

I'd say that's the only limit is

basically what the amount of data that

can fit in memory.

Okay.

All right. Were you guys able to run

this? Were you guys able to run the

linear regression ffit?

Okay, perfect.

Perfect. Do you Okay, great. Great.

So, we have a model and we can use it to

predict. Um, and so that's actually what

we're going to do next. If we go down

here, um we're going to have a function

that's going to um build a scatter plot

of our original test data.

Um so we're going to have our test data

here.

Um,

and we're going to then take our uh

we're going to take our training data

and plot we're going to use the uh this

data versus our sales predictions. So

you can see we're going to you this is

how by the way this is how you use the

scikitlearn model to predict. You have a

fit to train it and look at the function

you use to predict. It's literally just

called predict. That's how easy it is.

and you pass in your data, all your

features into this predict and it

generates a prediction for every row. So

every row in these features in this data

frame um will end up with a prediction

using our model. So what we're going to

do is plot our training uh features

against the predicted sales to see how

good of a fit that really was.

Okay. to see to see the regression fit.

Okay. And so there's the regression fit.

We have all of our test data here

plotted in the green. We have our blue,

which is our um we have our our blue,

which is our uh um training data line

that we built our model on. So that's a

pretty decent fit. Um and then our test

data is here. We just plotted in the

green scatter. But the thing I want you

to see is this prediction, right? We we

were able to generate some predictions

on that training um by running our

predict function with our model. Now

this model has been trained. So we've

already fit it and now we're using it to

predict, right? And so we're predicting

the sales and plotting that on the y

ais. So the sales are we're using the

predicted sales there which is our blue

line. So this is our line of best fit.

So this is our model prediction.

This is our model predictions. Right?

You can see it's a pretty decent uh

line, right? Pretty decent line of best

fit.

Of course, there's some error here. Like

there, you know, it's not perfect, but

it it does a decent job of being a best

fit line.

Okay.

So look how easy that was to

just to recap this to fit our model was

a linear regression.fit and of course

we're going to do more examples. So no

worries uh on that we're going to see

this many many many times throughout

this notebook. But we have linear

regression.fit to train it and then we

have linear regression.predict

to and we pass in our features and that

generates a predicted output.

Right? So what this is actually doing is

is computing this quantity.

We could do either.

We could do either. Um, so we could do,

so one thing we could do is plot uh, so

we could swap it out. We, we could do

either one. It doesn't, it's not a big

deal to do the training set. We could

do, so we could plot X test and then we

could plot linear regression X test.

So it's it's a similar line. Um it's

just different input features, but the

line is going to be the same. Just

different inputs,

but the coefficients are the same,

right? It's the same line. It's just we

generate different outputs.

So yeah, you could do either one.

This is This is honestly this is

probably better. I see what you're

saying. This is probably better because

this is the line of best fit through

this data. So that probably makes sense

to do to do predict on the test set.

Agreed on that. Probably makes about

most sense.

But you could do either one.

Yeah, I think that would be the most I

think that makes the most sense is for

it to be on the same one just to

validate. So like we could do we could

do training here and then train and

train just to see how that data lines

up. Really, what we're trying to do is

have our scattered data and then our

line of best fit on the same plot.

That's all we're trying to do, right?

So, yeah, I think I think they should be

the same.

I think that makes sense.

These values

or which values do you want to see?

Yeah, we could uh we could generate

those if we just do um let's go down

here. So the the line values

um are going to be uh the prediction. So

um the the

uh test

predictions

equals um

test predictions equals linear

regression.predict predict x test and

then we could uh we could print out our

test predictions.

Yeah. So we can see what those actual

values are on our uh on the test set.

Yeah.

Um we will do that. Yeah. So you thought

we were checking how well our data was

trained. We will do that. Yes, we

haven't learned how to evaluate this

yet. We're going to talk about that

coming up next. Yeah, we will do that.

We just haven't learned how to do proper

evaluation

of a regression model.

But yeah, it's something we're going to

talk about for sure

and see how to do in our code.

Okay.

All right. Any other uh questions on

this example?

Again, big takeaways

fit to train it and then predict to use

it.

Predict on the features to use the model

and make predictions with it.

So here is example. We we made all the

predictions. This these are all the

values that are on that line.

These are all our predictions. And

notice they this is a truly regression,

right? These are all floating point

values. Um so this is definitely a

regression, right?

Okay.

Uh, that's a good question. Um,

I'm not sure if there is

if there's like a verbose

there's not really no there's not really

a verbose. You can I mean you can look

at the source code if you really want to

see you can view the source code to see

um how it's done. I can tell you I mean

so generally linear regression is done

in two ways. Either you use a formula um

to to solve the optimization problem of

minimizing like this this uh distance

from the points to to the line. Um

or you use something called gradient

descent which is how a lot of these

things do it is they iterate through a

bunch of different iterations where they

update these weights according to um a

certain uh basically a gradient of the

the error function. The error function

in this case is the is the squared

distance from the line to the uh to to

the points.

So uh we can compute the gradient of

that and do um gradient descent. So if

you really want to look into it, I would

do some research on like linear

regression gradient descent.

Okay, linear regression gradient descent

to see how that's uh how that's being

done. Yeah, it it's it's a pretty simple

procedure. Um, again, you have the the

notion is that you want to minimize

minimize the loss or the error. Uh, in

this case, the loss is the square

distance. So, it's like um there's like

a it's a formula. It's like a sum of a

square distance from your prediction

um or your label sorry to your model

which is the beta 0 um plus beta 1 x1

plus beta 2 x2

etc like your model and then squared. So

this squared this is the squared

distance here and you're minimizing this

guy which is like a calculus problem.

You you find you basically find the this

is this is a I'm getting so far into the

weeds of this, but this is like a

parabola and you work your way No, no,

you're good. It's it's it's a good

question. Um you work your way down to

the minimum of it. Does that make sense?

Like you're working your way down here

and you do that through a descent

process, like a descent iteration.

Um

so

that's how these are found.

Um, but you don't see that happening in

the background. But if you look at the

source code, it I guarantee you it would

be it's either going to be this or

they're going to use the they're going

to use a a a matrix formula to basically

solve an equation um that involves this.

Basically, the derivative of this set

equal to zero and you find the minimum.

Either way, you're finding the minimum

of this.

Okay. But yeah, I don't think Psycharn

has like a uh maybe there's some type of

verbose flag you can look for.

I don't think they have that though. Not

that I've seen.

All right.

So I have uh an important um concept to

talk about next which is going to be uh

called overfitting and underfitting

um which is a really important concept

that's related to the training and test

data we just split apart to do

evaluation.

And um essentially the the issue with

machine learning is that it's not

perfect and it can struggle in different

ways. And the two ways that it primarily

struggles is going to be overfitting and

underfitting.

So overfitting is a situation where the

model basically memorizes the training

data so well that it's it fails to

generalize to new examples. So what we

see with overfitting is this exact sign

here where we have really good

performance on the training data. So

when so when we do that train test split

we see a really good accuracy or really

low error on the training data but it

does not perform anywhere near that on

that test data split. So what that means

is that the model is overfitting to the

training data. It's basically memorizing

it and it's not able to generalize very

well.

Now, why does that happen? It's usually

because the model is way too complex.

And that means generally you need to do

something to reduce the complexity.

Either you need to use a simpler model

or you need to use some type of

technique to mitigate overfitting. And

we're going to we're going to study some

of those techniques coming up in this

notebook. uh we might not get to it

today, but we're going to study

particularly what can we do to prevent

overfitting because overfitting is the

more common issue with machine learning

models. They tend to do so well at

learning from data that they pick up on

small details and patterns in the

training examples that they're exposed

to. They don't do a great job at

generalizing to new examples. They can

struggle with that.

So that's overfitting is struggling to

generalize to new examples, but you do

really well on your training data. So it

appears like you have a good model, but

it it's not able to go and make

predictions on test data very well,

which means we would not want to use

that model in the real world, right?

Because it's not able to generalize

outside of what it's already seen. And

that's not a good thing if we're trying

to use it for real world examples,

right?

So overfitting is a real issue. Um you

see it all the time. I've seen it many

many times in the real world, real

industry uh work that I've done.

Overfitting is a is a challenge for a

lot of machine learning models. And so

we need some techniques to overcome

overfitting. And we're going to study

some of those uh coming up shortly.

Um, one of the things that we can do,

one of the one of the things that we can

do to detect overfitting is exactly what

we just did, which is you split apart

your data into training and testing so

that you have a chance to do an

evaluation to see if you're even

overfitting in the first place. You want

to see that performance be consistent

from train to test, right? You want to

see consistency. What you don't want to

see is performance that drops off on the

test data. It's much worse. You don't

want to see that. That means that your

model is overfit uh to your training

data and it's not going to perform well

in the real world.

Okay. So, we're going to have a couple

ways to uh overcome that. Talk about

that. Um now, the opposite can actually

happen as well, which is called

underfitting. And underfitting

refers to the fact that a model is too

simple and it actually just performs

poorly across the board. So if we see

poor performance on the training and

testing data, that's a good signal that

the model's underfit and that means it's

too simple usually and you should try

using something more complex. Um, so the

best way to combat underfitting is to

use a more complex model. And as we go

through and learn about the models,

we're going to learn about which ones

are simple and which ones are complex.

So we're going to have a scale of kind

of complexity. And if you're

underfitting, you want to bump up to the

to a more complex model. If you're if

you're overfitting, one way of combating

that is to actually go down to something

more simple. Go the opposite way to

something simpler. So we need to learn

right now we've only learned linear

regression

but we will learn other models you know

in the future and we'll we'll talk about

uh their complexity and how they're

related to each other.

Okay, but these are two issues we see

just to draw that out again is if we

have a train test split where we have

7030 split let's say and we perform

really well over here but we go to apply

that model over here and it fails its

accuracy drops off significantly more

error that's that's definitely

overfitting which is not good

right and then underfitting is just not

performing well in either case so even

on the training data itself your your

accuracy is not very good. So you're not

really learning effectively. You're

underfitting your model. So that's

that's um underfitting case.

Okay.

All right. Now the issue is that it can

be very difficult to balance these two

and get it correct. That's what makes

machine learning a little bit

challenging is getting this balance

correct of simplicity and complexity. So

you don't want to be overly complex that

you overfit, but you don't want to be

overly simple that you underfit and

you're not able to learn effectively. So

there's a bit of a tradeoff there. And

this trade-off is typically known in the

community as bias variance trade-off. Um

in which case uh it's basically like a

complexity simplicity trade-off. It's

another word for that. Um,

and so, uh, it's it's thought that, um,

if you, uh, if you have very, um, if you

have a situation where you're able to

fit the training data very well, you

risk not being able to generalize. In

other words, you risk overfitting, and

it's hard to um, it's hard to combat

that in a way. Um, and um, on the

reverse side, if you have something

really simple, um, you risk not learning

enough. Even if you're trying to combat

that overfitting, you risk not learning

enough and your model just doesn't

perform as well as it could. So, there's

a bit of a trade-off there of trying to

find the right balance between something

complex enough to learn, but something

not overly complex that it's going to

not generalize to new data. That's the

challenge. Um, like I said, we are going

to have techniques to overcome this. So

luckily there are things to basically

overcome this trade-off and um and help

us along the way so that we don't

overfit. They basically prevent

overfitting

um and allow us to use complex enough

models um that that won't be overfit.

This is in the um this was in our uh

lesson 3.2 notebook. So you want to pull

that one back up. We were working on

Monday.

Um, and just to recap this a little bit,

remember we were building a linear

regression, I wanted to recap some of

the steps we took there, um, that we

will be doing over and over again. And

really the same kind of steps, uh, that

we do here, we'll do in a lot of our

model building. Pretty much all of our

model building um, that we do, whether

it's regression or classification,

doesn't really matter. um we'll still be

doing a lot of these steps which are um

remember first we split apart our data

into kind of a features and a label

uh x and y and the reason that's

important is because um the model

training uses the features and the label

um to help train the model, right? They

use those separately. Um so we want to

split those apart whenever we can. And

so we have usually uh it's a good

practice to call your features capital X

and your labels lowercase Y. And what we

do with that is remember we immediately

split that into what we call the

training in a test set. And the picture

we had for that was something like this

where we had about 70% of the data

we used to train the model against and

then the other 30% of the data we use to

test the model against. Meaning that we

build a model over here and we apply it

to this set over here um to make

predictions. And then the that's where

the supervised learning really comes

into play, right? is on this test set.

We already have the answers. We already

have the label. And so we can apply our

model to this to the features over here.

Predict uh what the the label should be

and compare that. We can get a a metric,

right, that compares how close we are in

our prediction to the actual values. Um

and that was some of our performance

metrics. I'll recap some of those that

kind of measure that distance away from

our predictions to what the actual label

is. Um, but remember we had this train

test split function which helps us split

apart our features and our labels into

these uh four sets of data. So we have

our training features, our testing

features and then our training labels

and our testing labels. So we have all

of those and um really these two guys

are going to be used to train the model.

That's why they're called underscore

train. They're going to be used to train

that model. And then the then we're

going to predict on these set of

features and then com use those

predictions to compare to this set of

labels, right? That's on the test test

set. Um and you notice here our test

size is set to 30%. Um, that's a pretty

standard number. Anywhere between like

20 to 30% is pretty standard. Um, we'll

typically use.3, but it could be 02.

Anywhere in between is fine.

Okay, so we had that. Hopefully that uh

we remember that from Monday.

So we had a train and a test set. And

then building the model was actually

really really easy. Once you have those

train and test sets, um, we just import

our model object. So from uh scikitlearn

sklearn

um linear model uh module from that

package we import the linear regression

model and then we do um linear

regression.fit

and we pass in our features and our

labels and this is again this is where

that supervised learning is really

coming into play because we're passing

in these labels.

That's really what makes this work,

right? We need those labels to help

guide the model to make those updates.

If you guys remember, the model is

something that looks like this.

So, this was a bunch of different

coefficients

um times the features,

however many we have. Um, and so these

labels are really taking the place of

this and they're helping us um make the

correct updates to these to these

coefficients or sometimes we call them

weights. Um, these B 0, B1, B2. Um, we

find out what the optimal one is to get

the best fit, right? To get the line of

best fit. Um that's what the model

training when we call this ffit ffit

that's really what it's doing in the

background is finding all those

coefficients right to end up with the

line of best fit that has the lowest

amount of error.

Okay so hopefully that makes sense.

That's just a dofit fit um to train our

models. And that's really going to be um

the case for

uh pretty much every single model that

we uh train with scikitlearn. It's

pretty much going to be a fit. We pass

in our training uh features and our

training labels.

Okay, so we had that and this was the

visualization of that where we had our

test points kind of scattered and we see

our line of best fit is the one that

goes through there with that minimal

error. That's that's the whole goal.

Pretty decent predictor.

Okay. And then we talked about

overfitting underfitting. So just to

recap this overfitting is the concept of

our model basically memorizing our

training data. It performs really well

on that training set but it is not able

to generalize outside of that. So it

performs poorly on the test set or data

that it's never seen before. Um and

that's overfitting. So the reason that

it overfits is generally the model is

too complex and it needs to be um it

needs to be simplified a bit. And one of

the things we're going to do today is

see a couple of ways we can alter the

linear regression model um if we are

overfitting to prevent overfitting. Um

so there's going to be ways to handle

this. Um and so we're going to explore

some of those today.

uh underfitting is kind of the reverse

of that. Remember, it's where the model

is not learning enough. So, the

performance is poor even on the training

data. It's not good on the test data

either. Um that is a sign that the model

is probably too simple and maybe we

should use something more complex like

go from a linear regression maybe to use

a polomial regression. Um or maybe use

an entirely different model altogether.

um if we're underfitting, our

performance is poor, it's a good signal

we should try something else. Um

okay,

so we talked about those

and one of the things we also talked

about was evaluations. If you guys

remember, we had different metrics that

we could compute to get a gauge of how

good our model is actually performing.

Um one of those was MSE, which is this

mean squared error function. Um so we

did this example during class last time

on Monday um where we uh were able to

generate the mean squared error. That's

one of our metrics. And we can see what

the mean squared error is on the

training set and see what it is on the

test set by um just passing in our um

training predictions and our training

labels, our test predictions and our

test labels. pass those into this mean

squared error function and it computes

the MSE and that's that's a helpful

function from the scikitlearn metrics

um package um or module I should say and

we'll be using that quite a bit to do

you know evaluation of of especially of

regression right mean squared error is

pretty is probably the most common uh

performance metric we can have and if

you guys remember what it's really doing

is measuring these distances So mean

squared error is kind of like the

average distance away from our our

points to the actual um to the

predictions which the predictions are

all on this line. Um so it's like

measuring on average how how much error

do we have on average right? Um, and the

idea is the closer to zero the better.

Generally means that the distance away

from our prediction to our points is

pretty low. The closer to zero it is.

Um, which is pretty desirable.

So a low MSE is kind of what we're

looking for. Um, closer to zero the

better. And so um if one model has if

one model has um a low lower MSE than

another, it's it's a better performing

model, right? It has less error.

Okay. And then we also looked at the R r

squared or sometimes known as R2 um

score. Um this is another metric that we

could use that measures the the

variability

um of uh the predictions and if our

model is capturing that variability um

well um and so R squar is has a range of

0 to one one is better that means the

model is capturing the the changes in in

the um output it um our predictions

follow along with those same changes um

so they're pretty close um so closer to

one would be a better score. So we have

those kind of metrics. So like on this

data um this would this would show that

this model was underfitting remember

because this

mean this MSE was bad and this MSE was

bad.

Um and what we should think of these in

the units of what our labels are. um

especially if we take the square root of

this the RMSSE that was another metric

we had um the square root of this is

actually in the exact units that we um

have for our labels. So uh in this

example this was the um this was the the

units or the sales versus the TV

products, right? Um and so this would

indicate that on average if we take the

square root of this um

and the square root of this um we have

uh

um we're on average about 11 sales units

off squared. So if we take the square

root of that um it's somewhere around 3

to four um somewhere in between three

and four units off. And this is as well.

Um, and because both of these are still

not close to zero. Um, this would be

under fit. And this shows that as well.

This isn't that close to one. It's

decent, but it's not um not that close

to one. So, we would say and performance

is poor on both training and test sets.

That's the key indicator of

underfitting. It's poor on both.

Yeah. Exactly. High MSE correlates to

underfitting. Yes. Yes. And it what's

key is it's high MSE on both on both the

training and the test sets.

If you have a high MSE on your test set

but a low MSE on your training set,

that's overfitting, right? Where it's

not generalizing from the training set

to the test data that it hasn't seen

before. That's overfitting. So the key

is high MSE on both sets.

All right. So we talked about that. Um

we did polomial regression last time. So

that was um doing

that was uh making a curved graph um by

transforming the features into polomial

features and then doing linear

regression with that. So you guys

remember from Monday we did this where

um we took our features and uh transform

them according to this polomial features

from scikitlearn. So we can go all the

way up to degree whatever degree we

want. So we put in four here, but

there's nothing special about four

really. This is just testing it out. Um

and we generate the the polomial

features and we can fit a linear

regression on those polomial features

and we get a slightly better model,

right? Um it fits the data a little bit

better than just a straight line. this

curved line with the polomial

features um performs a little bit better

and we could see that with the MSE right

we could evaluate the MSE of this um and

it would be lower

it would be lower than the curve line

and that's something we could do um we

would just have to pass in these test

predictions training predictions and

then the the test labels and training

labels and passes into the mean squared

error function and we could compute that

right wouldn't be hard to

All right. And then finally, where we

left off, um, you know, is on our

performance metrics. So, we talked about

mean squared error. That's that average

distance away from the labels to our

predictions. Um, and we take the square

root of that. It's it's basically

measuring the same thing, but it's the

square root of it is um more

interpretable because it's in the same

units as our label.

um mean absolute error is is the average

distance of the absolute value. So it's

not the squared distance formula like a

uklidian distance but it is a absolute

value. So it's a little bit um less

sensitive to outliers. They don't get

magnified as much. Um but it's not

typically used as much as a mean squared

error would be with regression. um we

talked about the last time because um

the distance formula or that distance is

actually what's used to train the model.

So it's a more natural um fit for a

performance metric for it.

All right. And then we had R square. We

just talked about that closer to zero

would be um worse. Closer to one would

be better. That means that the model

explains um all the variability in the

in the predictions. Uh it captures those

predictions um closely to the labels

um very well. So uh one would be better.

Closer to one would be better.

All right. So that's where we left off.

Um we're gonna pick up from there with

cross validation. Um, we've actually

already seen one method of cross

validation. So, we're going to study um

we're going to kind of recap that and

and then um talk about cross validation

in general um and look at some more

sophisticated techniques of it um coming

up next. But before I do that, any

questions about anything we've covered

um to this point in in the recap or

anything from Monday? Any questions on

that?

All right. So let's talk about uh cross

validation. Um now this term cross

validation refers to a technique that

evaluates performance. And what it does

is it divides our data into essentially

um training and test sets which we've

kind of already seen. And then we are

able to train a model on on the training

set, evaluate it on the test set, and

that's where that's where we get the

name cross validation because we're

crossing over our model from one batch

of data used to train it over to another

set of data used to validate those

predictions. Um, and there's actually

different ways to do cross validation.

So cross validation is a bit of an

umbrella term for multiple ways to do

that. We've already seen one way of

doing that um which I'm going to scroll

down to is um known as a hold out cross

validation. So that's um what we've been

doing so far. So this is just um

generating a train and a test set

train um split.

Um that's the that's what's known as the

hold out cross validation method. Um and

and this is exactly what we've been

doing so far, which is you split your

data into some type of split, usually

7030,

um of a train and test

and then you um train your model on this

section of data and then apply it to

this to evaluate performance. Right? So

that's that's what's known as the hold

out method. Um it is uh you know

relatively simple. It's pretty fast to

do. Um, but there are more robust ways

to try to divide up our data a little

bit uh more evenly. Instead of just

having one split, we can actually do

many splits, which is the idea of um the

next kind of cross validation I'll

cover. But hold out method is one that

we've already studied. It's the most

basic type of cross validation you can

have. Um so hold out this is the most

basic

and we we've already been we've already

been uh working with this type. Okay.

So we've we've already seen hold out

method. Let me uh explain to you a more

sophisticated method a little bit more

advanced of a cross validation um which

is known as Kfold cross validation. So

this is um going to be a little bit more

advanced of a technique but this is the

idea of kfold is that you take your data

set

and you split it into k number of what

are called splits or folds. So you take

your data and you let's say it was let's

say k equals 5. So we have five splits

here.

Okay. So let's say k equals 5. we have

five splits. So what we're going to do

is we're going to we're going to train

our model on K minus one of those folds.

So if K was five, we had five splits.

We're going to take our model and train

it on four out of five of those uh

splits. So let's say it's these four.

We train it on these four.

Okay. And then what we do is the one

split that's left over we will we will

test our model against that split. So

we'll test here.

Okay. Now, this sounds very similar to

the hold out method where we're doing a

train test split, but it's a little bit

this kful cross validation is a little

bit more sophisticated because we repeat

this process that I just mentioned over

and over for all combinations of the

splits. So then what we'll do, this is

just one trial that we'll do it again,

but this time we will pick um four

different splits.

So, this time we might pick,

let me do blue. This time we might pick

this one, this one,

um,

this one,

and this one.

And then those four we will train our

data on. And then we will test against

this one. Okay? And we'll do we'll

repeat this

repeat for all combos of the folds.

Okay. So we'll repeat that. So

essentially what we're doing is rotating

through. Every time we rotate through

one of the folds is going to be left out

as a test set. Now this is a little bit

more robust than just a train test

split, right? because we are exposing

our model to more of the data in in

doing this, right? Because we're going

to split it evenly into five or 10

splits. Those are pretty common um

number of folds to use. 10 or five. Um

those are the ones I've most commonly

seen. Um but we're going to by rotating

through which folds are being used for

training, which ones being left out. um

we are exposing our our model to more of

the data this way than just doing a

single train test split. Right? So now

what do we do with with the results is

every time we do this we we generate um

an MSE let's say or some type of

performance metric. So let's say we

generate an MSE from this guy,

we generate an MSE from this version and

we generate an MSE for all combos.

each combo we generate MSE and then what

we do is we average

the metrics

or the in this case uh if we use MSE we

would average those together. So every

time we do a fold combination and we

keep four of them for training, one for

test and we rotate through all those

combinations, we are going to generate

an MSE for every combination

then we're just going to average those

MSE's to get a final. So the final MSE

of cross val of this kffold.

So the final metric

is just the average of the uh

performance on all of the fold

combinations. Okay. So our final MSE, we

just average all those MSE's from all of

our combinations.

Okay.

Now, what's the advantage to doing this?

It's way more robust of a estimate of

the of the performance of the model

because we're exposing it to all

basically all of our data, right? We're

getting a sense of how it performs

across all those different folds. Um

rather than just doing a single train

test split, which is a bit it's basic,

it works, but it's a bit basic. Um so

this is more robust estimate of the

performance.

Now, what's the drawback to doing this

is that it's more intensive. So, if you

have a lot of data, this is going to be

pretty expensive to do because you're

going to have to especially you have a

high number of folds, right? You're

going to have to divide your data into k

number of folds and you're going to have

to do this over and over again. Um, and

if it's a large data set, it might take

your model a long time to train. It's

going to be a little bit more uh

computationally intense than if we just

did a train test split.

Okay, we just did a single like 7030

split. We only do that once. We only

train the model once, right? We train it

on the 70, apply it to the 30% test data

and evaluate performance that way. Um,

so we're only really using the model and

training the model once, but in this

kfold, we're going to do it um, you

know, k number of times essentially

or I should say one for every

combination that we have to work through

of of all the folds.

Okay.

All right. Does that make sense? Any any

questions on Kfold cross validation? So

K K K K K K K K K K K K K K K K K K K K

K K K K K K K K K K K K K K K K K K K K

K is an important uh number here. It

it's how many folds, how many splits do

you have? A typical value for K is going

to be somewhere like five or 10.

So 10 folds or five folds. Those are

pretty pretty standard

from what from what I've seen.

But does the does the concept make sense

or is there any questions on it on in

terms of um you're always going to leave

one fold out. You're going to split it

up into K number of folds. Always leave

one out. Train on the rest of it.

Evaluate on that one that gets left out

and then rotate those through. And

you're going to do that for every

combination and average all those

metrics.

And by the way, there's going to be an

easy function in scikitlearn that will

do this for us. So managing all these

combinations will be really easy. It's

actually just built into scikitlearn. So

we don't have to um we don't have to do

this all by hand. Okay, this will be in

scikitlearn. It'll handle doing all

these combinations of folds for us and

computing the average metric will be

really easy. So um

we don't have to worry about that. We're

going to see an example of this coming

up shortly.

All right, of kfold cross validation.

But this is a this is a really widely

used technique. And again like the

purpose you may be wondering like what's

the purpose ultimately of doing this?

It's to get a sense of if our model is

going to perform well on new data.

That's really what we want to know. like

is the model going to perform well when

I start to use it on new data that it's

never seen before and this kffold is a

decent indicator of that because we are

varying which data it sees across many

different folds right so it's a it's

kind of a good um proxy to exposing it

to different kinds of data each time and

seeing how it performs

right all right because we're working

our way through each one of the folds

there's always going to be one fold left

out. We're going to change which fold

gets left out each time. And um that's

sort of mimicking the idea of we're

going to apply our model to new data and

see how it performs. And it's it's new

data every fold.

um how we know which model is best suits

for which scenario because we have Yeah,

that's a good question. Um, so my we're

going to learn this as we go along

because we haven't covered all the

models yet, but generally the best

advice I can give on that is

you you generally want to start as

simple as you can get and then if it's

not performing well then work your way

up to something more complex.

So we are going to have models that are

simpler. We're going to have models that

are more complex. The rule of thumb is

to start with the most simple model that

works.

So you're usually going to have the same

ones that you're going to try in the

beginning. And linear regression is a

very simple model. It's usually the

first one you want to try for regression

because it's the simplest.

Um, and for classification, we're going

to have a similar like logistic

regression is the simplest kind of

classification model we could have. So

usually want to start with that and then

if it underfits like if we see it's

producing a lot of error then we work

our way up to a more sophisticated

model.

So um that's the way we that's the way

it should usually go is simple to

complex it based on their performance.

So we evaluate it and then we can repeat

the process. If it's not performing well

we can try something different that's

more complex if it's underfitting.

Uh this is a good question. Does a model

reset after training each k minus one

fold? Um yeah, it's essentially like a

blank model every time uh every fold. So

um we imagine like you have a brand you

have a fresh model every um k minus one

combination. Yes.

And the reason the reason it has to be

that way is because you don't want the

other folds influencing the model that

like on on the next combination. You

don't want the previous combination to

influence the results on the next one,

right? Um you want it to be a fresh

evaluation on every combination of

folds.

Okay.

All right. So, let me describe to you a

variation on what we just um talked

about with the K-fold. So, there's

another cross validation known as

stratified K-fold. And um this is the

same exact procedure as kfold except

that when we this is used for

classification.

Um so when we do classification

uh we want to make sure that the

different categories are going to be um

split amongst those folds in a

proportional way. So we don't what we

don't want to happen is um when we split

apart the data. So, let's say we have

let's say we're predicting um spam not

spam. What we don't want to have happen

when we do our splits is we don't want

to have all of the spams end up in one

fold and then every other fold has no

spam, no spam, no spam, no spam, right?

That's not very good. Um because if we

if we train against all these guys, we

have no shot at predicting spam when

they've never seen spam before. So

stratify kayfold is is used in

classification

and it's to um it's to make our splits

ensure that they have basically a

balanced number of categories for each

split. Um so that we don't end up with

certain splits with way more spams than

not spams. Um so we we do what's called

stratifying where we make sure the

proportions are balanced across each uh

split. So this is only really useful in

classification, not really necessary in

regression because we're predicting a

value. But if we were predicting a

category,

like in classification like fraud, not

fraud, we don't want to do the split and

have every single fraud example um by

bad luck in our shuffling and split end

up in one split and every other um every

other split has no examples of fraud.

Right? So we want to stratify this to

spread out those um frauds against all

the other splits. Um so uh again um

scikitlearn will take care of that for

you. Um but if you're doing

classification and you have an

imbalanced data set um you you really

want to make sure you stratify k-fold.

um imbalanced meaning that you have a a

um different number. Like if you're

doing fraud, not fraud, you have way

more not frauds than frauds. Um where

where that category is imbalanced,

you want to make sure it's balanced

across all your splits.

Um so this is this is useful in

classification only, not really

regression, which is what we're talking

about right now. Um but it's just a

variation on this that ensures when we

do those folds um the data is

distributed evenly amongst those folds

as much as we can. The labels are I

should say.

Okay. So that's stratified kfold. It's

the same same procedure once we have our

splits. It's the same where we do k

minus one of them. We train test on that

last fold um and then rotate through all

the folds and and average all the

metrics. the same exact procedure. It's

just the splitting itself um is going to

be balanced in a stratified kfold.

Okay. So, hold out we've already talked

about um is just doing a single train

test split. We've talked about that. One

more variation that is a bit of an

extreme version of K-fold. So it's

actually the same process as Kfold, but

it's an extreme version is if you set K

equal to the number of data points. So

you basically are um this is a really

really extreme kfold where you um

basically are training on all the data.

Um so you're training on all the data

except one point and then you test

against that one point. Um now why would

you ever do this? Um it's mainly so for

this reason here. It's to um maximize

the amount of training data that your

model gets exposed to because instead of

just doing instead of just doing five

splits

um which would be like

you know these four folds are going to

be used and then we um test against one

fold. um we're essentially going to use

99% of the data, right? One point is

going to be left out. 99% of the data

gets used to train. Um and then we're

always going to leave out one point. And

and the issue is we're actually going to

do that over and over and over again and

rotate that one point to cover the whole

data set. So, we're going to train on

99%, leave one that one point out,

and then rotate through every

combination of points until we've left

out every single point, and then average

all those together. Um, so this is a

this is an extreme kfold. Again, the

number of folds is actually equal to the

number of data points in this case. So,

we have every point is its own fold and

we train on everything but one. Test on

that one. This gets you the maximum size

of your training data because you're

basically gonna have every point but one

used in the training.

This gets you the maximum size. However,

it gets you the maximum uh expense

especially for large data sets. This is

going to be usually you're not going to

use this um especially for large data

sets because it's just too extreme. It's

going to take you a really long time to

work through every single point being

left out. um it's just going to take a

while to do.

So, for that reason, the leave one out

um that that's why it's called leave one

out because it's you're leaving one out

every single time. Um is rarely used. I

I have don't really see it used that

often, but it is an extreme version of

kful cross validation.

Okay. But rarely ever actually used. I

think the the ones that get used the

most are definitely the hold out method

with just a regular train test split. Um

and then uh the other one that gets used

quite a bit is is kfold

or stratified kfold if you're if you're

doing classification,

but certainly kfold in the in a

regression case.

Okay.

All right. Um we're going to do an

example with these guys. So we'll do

that next.

um with with the different cross

validation techniques. Um but any

questions on what they are doing

conceptually before we actually do the

code example.

Okay.

Very good.

All right. So, let's see some examples.

Um, let's go into our code and build a

model and do the different cross

validation techniques on it. Um, you're

going to see it's actually going to be

really easy to do and we it sounds

complex like doing the kfold and leaving

one out and testing. It sounds kind of

complex, but I promise you scikitlearn

makes it really easy to do. Um,

and so, uh, we won't need to do too much

besides just use the right, uh, tools

from scikitlearn. Uh, so we're going to

we're going to see that. Um, so here we

have some imports. The, um, primary, uh,

thing that's a little bit new for us is

going to be these, um, different kinds

of cross validation techniques. So we

have our kfold, we have our stratified

kfold, leave one out. Um, which are

those different cross validation

techniques. Um, these are going to be

used in combination with this cross val

score which is going to keep track of

the different um metrics and then

average them

uh while we do one of these um cross

validation techniques. So this guy gets

used in combination with one of these to

um as as we're going to see in the code

uh to average those metrics um doing the

different folds, right? Perform doing

performance against the different folds.

Okay. And then of course we need a model

using a linear regression. That's that's

the one we've studied so far. Um and

then we have just a regular metrics. If

we want to compute those um using maybe

just hold out, right? And hold out um

which which is just a regular train test

split um we could use these guys to

evaluate performance.

But in a more sophisticated kfold style

of cross validation, we're going to use

this to evaluate the the performance.

Okay, let's see.

So, we're going to be working with this

housing with ocean proximity data. Um,

you guys should have this one. Uh,

so you guys should have this one. So, if

you want to follow along and run it

yourself, um, you can load that one in.

Um, I want to make sure that I have it.

Let me pull that one in. So, it should

be this guy.

I'm going to load that in so I can make

sure I run it with you guys.

Um,

let me run this.

Do you guys have that data?

the housing with ocean proximity.

It's another it's another housing data

set. Um

but it it's a little bit different than

the ones we've seen before. It has a a

special feature for how close it is to

the ocean at different locations.

So it looks kind of like this. If we

load it in and do our head, which is

usually what we do, right? We can see um

we can see that it's got these features.

So it's got uh uh bedrooms, total rooms,

um it's got uh median age. Now this is

this is looks a little strange for total

rooms and um uh bedrooms and population

etc. But it's um

it's it's got those uh it's got those

because it's representing an entire

neighborhood. So it's an entire

neighborhood. And we're looking at this

um this is actually going to be our

label is this median house value for the

entire neighborhood. So what's that

median value uh in the neighborhood? And

this is the total number of bedrooms,

total number of rooms, um population,

households. So, how many houses are

there? Um, median income. And of course,

these are scaled. So, these are um

likely times, you know, uh thousands. Um

but um that's our data. We could

describe it.

So we can see the average age, average

median age. Um which sounds a little um

weird, but that's it's because again

this is the median of data within a

neighborhood. Um so the average of those

is about 28 or 29. Um we have

u

total bedrooms. The we can look at the

men. There's some data that only has

one. So, it's likely only one house in

there. Um, which is what this

represents. There's only one house. So,

there there is some neighborhood that

only has one house. Um, and we see the

median um we see the minimum uh median

house values there. And then the maximum

down here um is a pretty big number.

6,000 households is the largest that we

have in any any one of these

neighborhoods.

Okay. So, just a little bit of

description of the data.

Okay. So, then we can run.info. So, this

is um let me ask you guys, were you able

to load this? Were you able to run this?

If you're following along, were you able

to

load it and take a look at

Okay, great. Great.

Okay, so we're able to load that and

then look at head. Perfect. Um

Okay.

Um and then we run describe which gives

us that uh usual kind of statistical

description. Uh so we can see some

interesting stats about those.

What do you guys notice about the info?

Anything interesting that we see from

there?

Is there any missing data

any features that have missing data? Can

we see

object? Yeah, object type usually is

string. If it's an object type, that

usually means string. Python when we

read it into pandas it usually is just a

string.

So that that makes sense like we have

mostly numerical features but then we

have a this ocean proximity which is a

string.

Yeah. Total bedrooms has nles. That's

right. Because you can see here this

does not equal the number of uh rows

that we have. So this is the number of

rows which about 20,000 rows. That's a

good size data set, right? 20,000 rows.

That's decent. Um we're definitely

missing some data here for sure. Um we

could count how much we're missing

exactly by running this is NATO sum. Um

and so we see that total bedrooms is

missing about 200 uh 200 rows are

missing total bedroom uh value.

Okay. And then one thing I wanted to

look at is yes, this is a string. So

what remember what we can do with those?

That's a categorical.

So ocean proximity

is a categorical

string

feature.

So we can take a look at its value

counts, which is usually a good idea to

take a look and see what possible values

that feature could be. So if we look at

our

um what are we calling this? Housing

data.

housing data

ocean

proximity

value counts.

So, here's the different types that that

one can be. So, there's some

neighborhoods that are less than 1 hour

from the ocean. There's some that are

inland. There's some that are near the

ocean. There's some that are near a bay.

There's even five of them that are on an

island. So, these are the different

values of the ocean proximity. So,

remember, you can always do that. If you

see a string feature, you can always

take a look at what its um categories

are. And it looks like most things are

less than 1 hour from the ocean, but

it's kind of evenly distributed here. Um

otherwise

very few islands.

But as you can imagine like this feature

is probably going to be important for

determining um what the value is, right?

Probably going to be important.

Okay. So, um, we need to deal with these

NLES. If we're going to build a model,

right? So, um, this is all of our

typical data prep. If we want to build a

model, we're going to have to deal with

these NLES. What do you guys think we

should do with the NLES? What would you

what do you think for total bedrooms?

What do you think is a good strategy to

do? Keep in mind, we have 20,000 points,

20,000 rows I should say, and about 200

of them are null.

Right. So about 200 are null. Um so what

do you what do you guys think would be

like a good strategy to deal with those

NLES in that case?

average. We can't ignore it because we

can't ignore that column.

We can't ignore the whole column. So,

something needs to go there.

Probably don't want to make it zero.

I think average is a decent average is a

decent idea. Probably don't want to make

it zero because um that would indicate

that there's no bedrooms and yet we

still have a bunch of total rooms. So it

probably doesn't make sense to do zero.

Average, I think average could be a

decent one.

Now in this example, what we're actually

going to do is we're

rows.

We're actually going to drop the rows al

together. Now, why are we doing that?

It's because we have so much data and

only 200 of them are null.

Okay, only 200 of them are null. So,

we're actually just going to drop the

rows. Now, that's a choice.

Um, that's a choice, right? Is that we

could fill in with the average like you

guys are suggesting. What we're actually

going to do is just drop the rows. It it

makes up less. It makes up about 1% of

the whole data. So it's not that much of

it is missing. We can drop those rows.

So that's actually what we're going to

do here is we remove all the roles with

the NLES by doing drop NA. So this just

drops them. So those rows are cut out.

Um, it's arguable that we could replace

it's arguable that we could just replace

it with something and I think you guys

have good thoughts which is the average

a default

um assume total bedrooms. We could we

could try that. Yeah.

Assign a value based on comparable home

value. Yes, you could do that too.

That's a good strategy is to look at the

other rows that are similar to it and

fill in a value. That's absolutely fair.

Um, in this example, we're actually just

going to drop those rows,

but I think that's totally um totally

valid.

This is a choice.

We could fill NA with different values

such as the average

total bedrooms

um derive a value etc. So we could

derive something which I think Brent you

have a good suggestion that's a good

suggestion. Um we could derive something

like that uh and fill in the blank and

that's I think that's totally valid. Um,

we could take the average of the um

bedrooms. Uh, I meant total rooms here.

Sorry, total rooms. Um, we could fill in

we could fill it in with the total rooms

for that category um or for that row.

Um, many options. In this case, we're

actually just going to drop those rows

because they make up such a small

percentage relative to the 20,000 rows

that we have. It's about 1%. Right? 200

rows is about 1% of 20,000.

So, we're just going to drop them. But

that's a choice. We don't have to drop

them. We could fill in with something.

Um, and if we did that, we would use

fill NA rather than drop NA, right?

Uh after dropping the rows, how many? So

it's just so after we drop the rows, um

after we drop the rows, it's just going

to be we still have all our other rows

are intact, right? So if we look at this

now,

we now have um slightly uh slightly less

entries.

So now we have this this many um rather

than rather than this many,

right? We dropped those 200

But they're all filled in. Yeah, they're

So all the other columns are still

filled in. We're just we're we're

cutting out the whole row. So if you

think about our data set, um we have all

these rows and all these columns. What

we're doing is like if there's a null

here, we're just we're just getting rid

of that whole row, right? And so we

still have all the other rows intact.

Uh, we can drop them because we have a

good sample size. Yes,

that's exactly right, Ronald. Yep, we

can drop them because we have we have

20,000 rows and only 200 are missing

values. So, that's totally fine.

Uh, drop a removes all rows that has any

null. Yes, that's true. It it will go

ahead and just drop any row where

there's any null, no matter what column

it's in. Yes,

index. Yeah, the index is not getting

reset. Um, that's true. So, um, what we

what you can always do is you can reset

the index. So, um, if you want to, it's

optional. We we're not really going to

use the index for anything that

important, right? But what we could do

is, uh, reset index.

Uh,

we could do that, right? Which will

reset it.

So now now it gets reset.

But um let me actually I don't I don't

really want to do that. I'm going to

reset this.

Um,

yeah, we could do that.

Okay.

So now importantly there should be uh no

missing data of this of this new one

where we've dropped NAS. Right. So now

this is good. If you now the reason we

had to do this is because if we try to

build a linear regression and we have

NLES in there. Um the the issue is like

how do you build a model where you have

something like this

and these are null? Like what do how do

you multiply a number by a null?

Um we can't really do that, right?

we can't really do that. So, um,

so therefore, uh, we need to get rid of

NLES like the the null is not really

going to work in there. So, uh, we need

to get rid of them for linear regression

to to really have a chance to work,

right? To train it and be able to use

it.

You got to get rid of those nles.

All right,

any questions so far? So, we haven't

done any modeling yet. We're doing some

We're doing some data preparation before

we get to the modeling. And we haven't

done any cross validation yet. We

haven't set that up. We're just doing

our data preparation before we get to

the modeling. Right? So, we've dropped

some NAS. We've checked it. Um, we're

going to do one more prep step, which is

to um change that ocean proximity

feature into something numerical because

again, how do you build a model where

you're inserting a string into those

like beta 1, beta 2, beta 3 times of

features? You can't really do that when

it's a string. Um, so what we're going

to do, I'm going to get rid of this

because I don't think we really need

that. um is we are going to uh run this

get dummies function which is our um our

get dummies function is our usual one to

uh our git dummies one is our usual one

to um

uh get our one hot encoding.

So this is our uh one hot encoding here.

We now are going to have data that's

like this, right? So we have ocean. So

So by the way, this prefix

um this prefix is OP, which which is

short for ocean proximity, right? So we

have ocean proximity uh less than 1 hour

from the ocean, ocean proximity inland,

ocean proximity island, near bay, near

ocean. So these first five rows are near

the bay. Um so they have a one there and

a zero in the other spots. So this is

good. This one hot encodes that feature

into these numerical uh values,

right?

Were you guys able to run that one? they

get dummies.

So the reason that Yeah, that's a great

question. How did it go ocean proximity?

It's because um that is the only uh

string feature we have. That's the only

one we have. So it it's going to look

for any non-numericals and one hot

encode those however many however many

there are. So whatever objects we have

which are strings, it's going to

automatically oneh hot encode those.

Yeah, we could have Right. We could have

went here and did Right. We could have

done ocean

proximity,

but we only have one of those features.

So it's just going to do that to the

whole data frame

uh on that one feature. So what we're

going to do is um go ahead and split it

into an x and a y um which the x is

always what includes our features. The y

is what we are trying to predict which

is the label. Now, um, in order to

separate those out, what we're going to

do is assign X to be the variable that

is, um, our data frame minus this median

house value column. So what this is

doing is um uh it's not permanently

dropping because we're not uh dropping

it in place but it is returning us a

copy of the data frame with the median

house value column left out right it's

dropped. So this is this is uh something

we want to do because that will the rest

of it will contain our features right.

So, um this will temporarily or I should

say return a copy of the DF with um

median house value

dropped,

right? Median house value dropped. Um so

we go ahead and drop that one. Uh now

remember it's not permanent. It's just

giving us uh the remainder of it which

is this housing data. dropping this and

it's assigning that to X and then we're

taking the actual median house value

column from the original data and

assigning that to Y. So this is going to

be our labels,

right? So this is what we are trying to

predict.

Okay, so that is our Y and that's always

how it is. X is our features, Y is our

labels. Um hopefully that makes sense.

What this is doing is this is going to

get rid of that label column and

everything else will be our features and

then this will get rid of this will just

assign the label column to Y.

All right. And then what we can do is

pass X and Y into our train test split

function and this will generate the hold

out set. So if we want to do the hold

out cross validation this is how we

would do it is we would split the data

into X train X test Y train Y test um

using train test split. So this is what

we did last time. This would be this

would be for hold out cross validation

right where we are uh uh just have that

one one set for testing one set for uh

one set for training one test one set

for testing I should say right so this

is pretty standard train test split um

we pass in that x we pass in the y we

use a 30% test size which pretty

standard

and random state so that we get the

consistent shuffling if we were to run

this multiple times. Um we we get that

uh consistent randomization.

Okay,

so we have that and so now our X train

is a percentage um of the data frame of

the 20,000 uh rows and the X test is uh

30% of that. So it's only about 6,000

rows, which is what um the shape of that

is.

Yeah. X. So X is our features. So we're

we're putting all of our data in that is

our features into X. And so the the um

most efficient way of doing that is um

the most efficient way of doing that is

to

uh just take our data and drop the

median house value column because that's

our label column. So we just remove

that. The rest of the data is our

features. So that's what that's what

this X is, right? It's all of our

feature data. All of our columns that is

not the label column essentially is what

that's doing. And then Y is our label

column from our original data,

right? Y is our label column. And so

this this will um contain all of our

labels which is the median house value.

X X contains every column but the one

we're going to so we we ultimately

decide that but X contains um X is

everything that is not our dependent

variable which is what we're predicting.

So we're removing what we are trying to

predict from X. X should be everything

else. That's always how it's going to

be. X is X is always going to be all of

those independent variables that we're

using to predict the median house value.

So we are going to predict the median

house value. We need to remove it from

X.

So we're we're taking everything but

that column.

So it's the whole data frame. It's the

whole data frame minus this one column

with just the dependent variable. Right.

Exactly right. Removing the dependent

variable and keeping all the

independence. That's exactly right.

Exactly right. So think about it in

terms of the model. Let's go back to the

features. Right. Think about it in terms

of the model. We are trying to predict

this this value. We're building a model

to try to predict this. So we are going

to make sure x is everything but this

right. So this is actually just y.

That's our label. That's our dependent

variable. Right? That's y. Everything

else is belongs to x. Everything else

belongs to x including all of these.

Right? We choose this one to be y

because we're building a model to

predict that. That's our label.

All right. So, we have our we use X and

Y to do our train test split. So, we

have our our training features and our

test features and then our training

label and test labels here. Um, pretty

standard there.

Um, okay. So, this is what's new is if

we want to do k-fold uh validation, what

we're going to do is create a kfold

object. So, we have this kfold from

scikitlearn that we already imported. we

are going to create a kfold um where we

are going to specify how many folds we

want. So that is the in uh inslits

parameter as this says um this is going

to be uh uh in this case we're going to

do 10 folds. That's pretty standard. So

I think the typical number of folds that

I've seen and I've worked with in my in

my career is usually five or 10.

Five or 10 folds is the standard.

Okay. So, we're doing 10 folds in this

case and we're setting a random state

because we're going to do shuffling. So,

in order to produce those folds, we're

going to shuffle the data first and then

split it into five folds, right? So,

this this kffold object is going to

manage creating these splits for us,

right? These even splits. I know I I

didn't draw it even, but um it's going

to manage these five folds for us and

it's going to shuffle the data and

assign them to these different folds and

we're and then what we're going to do is

use those to do our training.

We're going to execute the cross

validation using this kfold object.

Okay, so we create the kfold

um we initialize our model as well. So,

of course, in order to train something

uh in the K-folds, we're going to need a

model. In this case, we're using linear

regression, right? Which is which is the

model we've been studying so far. So,

you have a linear regression. Um now,

look how easy it's going to be in order

to execute cross validation. All we need

to do is um all we need to do is create

a cross file score function

um or I should say use the cross file

score function from scikitlearn. So we

use that with the model we want to

train. So our model goes first. So

that's the linear regression object.

Then our data. So our extra our features

and our label for our training.

And then um let me skip over this for a

second. I'll explain what this is in a

second. Um but then we are using uh the

cross validation technique is our

K-fold. So this is where our K-fold

object goes in the CV parameter which is

cross validation. So what cross

validation strategy are you using? We're

using Kfold and the K-fold we're using

is this one we defined up here KF. So

we're putting that right here for this.

And then um in jobs um allows us to

parallelize this. So if we set it to

negative one that's the that that's the

default um it will do it will actually

train across the different combinations

in parallel um which speeds it up. So

you want to you want to keep this to

negative one if you can. So um now let

me describe the scoring. So what this

means is we put in our metric here. Um

and so you can put mean absolute error,

you can put in mean squared error. Um

those are the two that we can use. And

um the reason we it has a negative in

front of it is because we want to find

the one that has the lowest score.

That's going to be our best model is the

one that has the lowest score. So, we

take the absolute value.

I'm sorry. We take the abs the the the

metric and we take the negative of it.

Um because the highest scoring one is

going to be the closest to zero. Um so

it's just a we use the we use the

negative of the of the metric. Um

because on the number line like the the

highest um scoring one should be the

least um or I should say the maximum

negative that we can get. That's going

to be closest to zero. So if here's

zero, this will be like -1 is better

than -10. Right? So something that

scores um the maximum negative uh

absolute error would be closest to zero.

And something that has more is going to

be on this side.

So this is only the reason we need this

is only just to keep track of the scores

of each individual um fold. Okay.

So the one so the reason we can do that

is at the end we can kind of see which

which combination performed the best. um

it's going to be the one that has the

highest uh highest value of the negative

which is closest to zero.

That's just a convention.

Yeah, it's just because um it's because

the cross validation is looking to

maximize the metric. So whatever has the

best score

um whatever has the best score is

considered the best uh performance. Um

but we are using uh something where

lower is better. So we we take the

negative and like the the highest

negative would be closest to zero,

right? The highest negative is going to

be closest to zero.

So that so it's it's just because like

we want the lower score to be the best.

The lowest score should be the best.

So we take the negative of it. Um and so

something that is more negative is going

to be worse. Yeah, that's the reason.

So something that's down this way is

going to be worse.

Okay. So it runs this

and what you can see is if we actually

print this out, if we print out our

k-fold scores, what we should get is 10

different scores.

And you can see um we have 10 different

uh scores here, which are all negative

because we're taking the negative of the

absolute of the mean absolute error. Um

so what we would be looking for here is

um we want to take the average of these

scores but take the absolute value of

them to get the best performance. So

this is capturing like this is the score

on the first fold combination. This is

the score on the second fold

combination. This is the score on the

third fold combination and on and on and

on. And these are the absolute errors.

Okay, these are the absolute errors. Um,

so if we take a look at computing the uh

average, which by the way, we don't need

this import because we're using the

numpy average. So that's fine. Um, we

can take the absolute value of those um

and take a look at the average MSE

or sorry MAE. Now I want you to think

about this this uh average performance.

So this is our performance right here on

the cross validation.

This is our average

M AE across all of our fold

combinations. So that's a that's an

indicator of our performance, right? Um

for the cross validation.

Now what are the units of our original

uh the original median value? They're

already in the thousands, right? So if

we go to that feature, they're already

in these hundreds of thousands. So this

is not a very good error. It's it's kind

of high, right? Because it's in this is

49,000.

Um that's that's how far away we are in

absolute value on average is 49,000 um

dollars on the median value. That's not

very good. So this score

this score is

um not very good. So this model is not

performing that well and we can see that

by comparing this error to our actual uh

data. So this is right around 50,000

and our median uh house values are in

the hundreds of thousands. So on average

we're 50,000 off when we make a

prediction. That's a significant amount,

right? That's a significant amount on

average um when our when our data is in

about the hundreds of thousands here.

So we are um we have a significant

amount of error 50,000 relative to the h

to our units that our our data is in.

Right? Um so this score is not very

good. Um

and so we see that from the cross

validation. So look how easy the cross

valid is. Again we just do cross file

score. We put in our model. We put in

our data. We put in our cross validation

uh strategy here which is kfold. And we

can generate these metrics across all

the fold combinations. So it's this

function is taking care of rotating

those and doing every combo with just

the 10 different combinations here of

the of the folds.

10 different instances where you have

you know 10 different folds are the ones

that are left out for evaluation.

Um so it's managing that for us using

this data right using this training data

here. Um and we uh we generate these um

generate these scores.

Okay. So that's kf fold. It's not hard

to do. All you have to do is um just use

a cross file score. And we could change

this to mean squared error. That's you

know we could do that too. That'd be

pretty easy. Um, so that'd be no issue.

We just happen to be using the absolute

error here. Of course, we could use

squared error.

Were you guys able to get this to run?

K-fold scores.

It produces an array of 10 10 different

scores, which should make sense because

those are these are the um we're

splitting our data into 10 different

folds,

right?

10 different folds. than leaving one out

to do our evaluation on. So the one that

gets left out every time is what's

producing these scores. So it's 10

different ones get left out when we

rotate through all the combinations.

And so we average these scores

and we get this amount. We get about

50,000 in error on average.

Um, what do you think would be what do

you think would be acceptable? So, if

our if we're predicting the price, like

if we're a real estate agent and we're

predicting these prices and they

typically are

Yeah, close to zero would be great.

That'd be fantastic. Closer to zero

would be better. The average is um

206,000.

So 50,000 is a decent percentage of

that. Um so you know you can compute it

as a percentage right. So 50,000 is a

decent percentage of that. Um probably

you want this to be less than 20,000

would be about 10% error. 20,000

right? So maybe like 30,000 somewhere in

there.

Yeah. 10% would be 5% error. 10,000

would be 5% error. That's true. That's

true. So that would be that would be

much better. So being closer to zero,

like the smaller the better, of course.

Of course. Um but yeah, I would say an

acceptable percentage of error is

probably 20%.

Probably 20%, which would be um like

40,000 or less would probably be

acceptable.

Usually when we usually when you build

models um 80% accuracy is usually uh

considered decent.

Usually considered decent

80%. So I'd say 40,000 or less would be

kind of ideal.

Does that make sense

to answer the question?

That's a good question. What value is

acceptable? I think probably less than

40,000 would be ideal. That's right

around 20% error.

All right, so that's K-fold. Um let's do

just a regular hold out now. So this is

just using our training and test data.

Um doing model.fit and calculating an

MSE on the test data. So this is this is

just the um hold out strategy here where

we just have um this is less robust but

it's a lot quicker to do and easier to

set up. Right? So um this is using the

hold out strategy. So just a regular

um train test split.

Are we going to rebuild the model? No,

not necessarily. There's some things we

could do most likely. And like one thing

we did not do was scale our features.

Remember I said that's a pretty

important thing to do is to scale our

features. We did not do that. So that

would be an enhancement to this that

we're going to So I I actually do think

we'll do that later. Yes. So I think we

will actually do that now that I'm

thinking about it. Yes. One of the

things we can do is scale these features

using like a minmax scaler or a standard

scaler. that's actually going to help us

um that's going to help us do better

predictions.

So that that's one thing we could do. Um

but yeah, we will we'll try to see if we

can get better.

It should help it. Yeah, usually you

want to scale you want to scale the

data. That's something we didn't do in

our preparation step. We did a lot of

the things we should do. We removed nles

and we did one hot encoding to the

proximity feature like this one. Um

those are good to do but we didn't scale

any of these other we didn't scale any

of the features right we didn't scale

any of them. Um it you it will have an

effect. It usually when we scale it

it'll be a better model.

It'll it'll learn a little bit better if

we can scale the data. Um so that way

like these

um like ages aren't you know drastically

different than like in scale than total

bedrooms or income

uh those kind of things. So we usually

want these to be in a similar scale

range.

So we'll we will I think we'll scale

them coming up in a bit and it should

help the model.

We've talked about that before, right?

Scaling usually is a good idea to do

when you're prepping your data for

modeling.

No, you want to you want to scale your

test data as well. You're going to do

both. You're going to scale your

training data. You're going to scale it.

So that's actually a good point you

bring up is any transformations you do

on your training to build your model,

you should also do on your test set so

you get an applesto apples comparison.

You should always do the same

transformations.

Yes. Would scaling data impact K? Yeah,

it could. It could make it better. It

could uh Yeah, it should impact it. We

should get a better model. So, when we

do the different folds, we'll get

different we'll get better scores. Yeah,

it it will impact

uh yeah, if they're so that's a good

point. If they're going to use our

model, then yes, they have to scale the

data as well. If they're going to if we

build the model on the assumption that

the input is scaled, then yes, they have

to also scale their data when they're

using it with our model. That's true.

I mean, not really. I'll show you why.

There's something that's actually going

to make it easier um that that will

automate doing the scaling for them. So,

they don't they don't have to do the

scaling manually. it'll just it'll

happen automatically when they use the

model. I'm going to show you something

that's going to automate that which is

going to be called a pipeline.

So that part will be automated and they

won't have to do that. So it won't be

heavy on the user. No, in theory it is,

but

has a really helpful tool to make it

easy to do that. So I'm going to I'm

going to show us that um later on in the

notebook.

No, the data data is not for a single

house. It's for like a neighborhood. So

there's a certain number of households

in the neighborhood. And this is the

we're predicting the median house value

of that neighborhood.

Yeah. So there's a there's certain

number of households. There's there's

like an a median income, a population,

certain number of people that live

there. Um proximity generally of where

that location is. It also has a latitude

and longitude.

So,

and a median age in that neighborhood.

So, yeah, it's not just a single house.

Okay, let's go back to this was the hold

out strategy. So, this is a lot simpler.

This is just model.fit, right? This is

just model.fit on the training uh data.

And then we um can predict on the test

features and generate test predictions.

And then we can compute our error on

those um we can compute our error

amongst the test predictions and our

test uh label. So that's our useful mean

squared error function, right? To to

compute the MSE. Um let's see what the

MSE is. So MSE is right here.

Um now what we could do is we can take

the MSE

and we can take the square root of it.

So let's actually do that. Let's um do

MP. Square root of the

um test

MSE

and we get um 67 we get 67,000.

So that's pretty high on this. So when

we just now look at the difference of

that, right? When we just do a train

test split,

um

when we just do a train test split, we

get a worse score because it's not as

it's not as robust, right? We're not

showing that to many of the other uh

folds. So, we get a lot more error this

way on the test data.

So, this is um actually worse

performance just doing the train test

split.

This is a really higher.

Yeah, we can. We can. I'm going to I'm

going to show us how to how the scaling

will be done automatically. Yes, we can.

Um there's there's a really easy tool to

do that will scale it automatically.

It's going to be later in this notebook.

I'll show us it.

All right. So, just to recap this, this

is fitting the model.

This is fitting the model. This is

making the predictions, right?

Model.predict.

So, this is making the predictions. And

then this is calculating the error, the

mean squared error, which is looking at

our test labels versus our test

predictions, right? And this is

computing the distance, the average

distance away from these values to these

values,

right?

And then we can also compute the R squar

R R squar and we see that it's not a

very good R squar 65 uh is not a very

great model

um because it closer to one would be

better. So this is still this is not

very good.

We know that we knew that from the cross

file score but this is just doing um

this is just doing a hold out uh where

we do a train and test split. Right? So

it's a little bit simpler but it's not

quite as robust. Um,

it's not quite as robust as the cross

valve, but it works. Um, it's, you know,

we can do hold out. Um,

we can do hold out, uh, to to quickly

evaluate a model and see if we need to

make any adjustments.

It's a little bit quicker to run.

Okay. And any questions on it? Does it

make sense what we're doing here?

Model.fit fit to train it predict to get

our predictions. Um this is pretty

standard, right? To train is the

model.fit and then to use the model to

predict we predict on the test features.

Um so this is passing on on all of our

features into this model to generate

predictions for every row. That's

something I also want to point out that

may be a little bit confusing is this is

a data frame. So we're passing in a

bunch of rows of features with columns,

right? So um we're passing in a bunch of

data that looks like this. And what

we're doing is essentially making a

prediction for every row. So this will

generate a prediction. This row will

generate a prediction. This row will

generate a prediction and on and on and

on. So this this predict will predict

for every row. And so we end up with

this collection of predictions here for

each row. and we're comparing those to

the labels that we have for those rows

from our from our supervised learning,

right? From our data set. So that's

truly supervised learning, right? We

have the examples and we're comparing

those to what our model is predicting to

to get our performance.

All right.

So let's uh let's try the other just so

you can see it. The leave one out. Now

the leave one out cross validation is

going to actually work the same way

where we put in the leave one out um

strategy inside of the cross file score.

Now here we don't need to specify how

many folds there are because we know how

many they're going to be. It's going to

be the number of data points, right? So

which is actually going to be quite

large because there's 20,000 rows. So

this is going to be extremely

uh extremely um intensive because we are

doing um you know 20,000 examples and

leaving one example out to be our

validation and then um doing that across

every 20,000 uh examples.

So we could do it though just to see how

it works. Um we have this again leave

one out. We generate our crossfile score

from our model our data and then same

scoring that we had before and but this

time we change our cross file to be

instead of our kfold object we have our

leave one out object which is this

um and then we could run this. We can

compute our average uh across the all

the folds. Now this is going to be a lot

bigger of an array. It's going to be a

20,000 size array and we're going to

compute the average across it.

So, let's do that. It's going to take a

moment because there's lots. So, if you

notice it when you run, it's going to

take a little bit of time to run because

it's running across all 20,000 examples

and leaving one out. So, you have 20,000

and then one left out to uh test

against. So, it's quite intensive. You

can see it's taking a lot more time.

It's still running. It's taking a while.

Okay, just let that run. Still running.

So, if you guys try running this, it's

going to take a little bit of time.

Hopefully, that makes sense why it's

taking so long, right? It's because it's

instead of doing 10 folds, it's it's

putting every data point but one is the

training set and then iterating through

all 20,000 points.

This takes a while to do.

Let's see what our

RAM our memory is a little increased.

Okay,

still running. That's okay. I'll let it

run.

Come back when it's finished.

Yeah, exactly. This is a this is for

this is giving us a performance

evaluation. This is like the average

error across all of our uh different

folds. Um now this is the extreme case

where we have the number of folds equals

the number of points.

Right? So it's an extreme case but yes

it's just like kfold. It's giving us

that performance estimate.

Okay. It's about the same. Right. This

is still around 50,000.

Not much difference, right? Still right

around there. But look how much longer

it took. That took 2 minutes to run. The

other one was pretty instant, right? So

this this took about 2 minutes to run.

So um definitely uh

yeah, definitely don't want to run this

uh too often. I think that it's

generally preferred to do k-fold. If

you're going to do cross validation,

generally want to do k-fold or just the

regular hold out train test split. Uh

generally better than doing leave one

out. It's just going to take too long

and um it results in about the same kind

of score as the kfold.

Okay,

any questions about um the cross

validation that we just did.

Okay,

good. And as it says here that the

stratified kfold is usually used for

classification. Again, we're not doing

classification yet. That's in going to

be in lesson four. So, we don't need to

worry too much about that. Just for

regression, um regular k-fold is

preferred, right? Because we don't need

to um worry about distributing

categories amongst our folds uh in any

regression problems.

And as we see the error is kind of high.

Um there's going to be some things we

can do to improve that which will be uh

later on we'll learn about some more

advanced models. This signals that the

performance is bad. We probably need a

more complex model. Um one thing we

could try before we try a complex model

is to do scaling. We will try to do

scaling. I'm going to show us how we can

do that coming up um in a in a nice

streamlined fashion. Um, but uh outside

of that, if we still had bad

performance, we would likely need to use

a more advanced model. And we'll learn

about more advanced models uh in the

next lesson. And what's great is some of

those advanced models can actually be

used for regression. So they have

variations that can be used for both

classification and regression, which is

pretty cool. So I'll point those out

when we get to them. Um, okay.

So what I want to talk about now is a

way we can combat overfitting. So if we

have overfitting which remember that is

the case where the uh the we see good

performance on the training data but

then um it doesn't generalize over to

the test data. We get poor performance

on the test data. Um there's there's a

drop off there. Um that would signal

overfitting.

overfitting

and one way of um combating overfitting

is to do something called regularization

which we're going to talk about next. So

the key idea in regularization

is to

change our uh the change the way we

train. Essentially, what we're going to

do is modify our training

uh error function or sometimes called

the objective function or loss function.

We're going to change that to add a

penalty to penalize excessive complex

complexity. Essentially the the way that

we're going to penalize is by making

sure the size of the coefficients

doesn't grow too much which should

mitigate overfitting because remember in

linear regression what we are learning

are the coefficients right we're

learning the beta 0 the beta 1 the beta

2 and on and on however many betas there

are beta n we're learning all of those

guys um through the regression error

function we're trying to minimize that

error function. That's how it trains. We

talked about that on Monday.

Um so what we're going to do is um

basically penalize the these guys

growing too big and making sure we kind

of keep them small so that no one

coefficient has a dominant uh effect on

the model. And this should help with

overfitting and complexity. It should

make the model simpler because all the

coefficients are going to be encouraged

to be smaller. They're not going to grow

too big. Um, and this this has the

effect of making the model so basically

make the model simpler.

Make the model simpler is what these

regularization techniques are

essentially trying to achieve is is

remove complexity, make them a little

bit simpler, make these coefficients

smaller so that you can generalize a bit

better and and prevent overfitting. So

we want to prevent

uh overfitting,

right, is what we want to do. Um so

there's going to be a penalty and I'll

show you where that penalty gets added

and kind of what it looks like.

Um but uh to control the level of that

penalty we are actually going to

introduce another parameter to our model

um called alpha.

Alpha is going to scale the penalty. So

if alpha is really high that imposes a

stronger penalty on the coefficients um

which will make the model a lot simpler.

So the higher the alpha the simpler the

model we will get and we the the risk

with that is we actually underfit. So if

alpha is too big we may underfit the

training data

um a bit too much because it will make

the model way too simple. Um and again

I'll show you what this means

mathematically in a moment. Um but on

the other hand if we have a lower alpha

this will have a lower penalty. it's a

weaker penalty term and that'll lead to

a model that is um a bit more complex.

Um which could um risk some level of

overfitting. Um so there's so there's

still the risk of overfitting if you

have a low alpha. And of course if alpha

goes all the way to zero there's no

penalty at all. So you're back to your

original linear regression um which

could risk a lot of overfitting.

Right? So you you generally want to pick

an alpha um effectively and actually

we're going to see h what's the best way

to pick alpha. Um we're actually going

to learn how to do that. I'm going to

show us how doing some tuning techniques

to pick what alpha should be. Um but um

a a pretty industry standard alpha that

most people default to is alpha equals

to one. So just just one which signals

that there should be some penalty. we

just have alpha equal to one is a

standard penalty. We don't want it to be

too high. We don't want it to be too

low. Like we don't want it to be a

fraction. Um but a penalty of one is

usually uh good enough.

Okay, I'm going to show you where that

comes into play in a moment.

Um but the whole purpose of doing this

is to mitigate overfitting, right? Um

that's what and and doing this penalty

is is called regularization. So adding

so going beyond just regular linear

regression adding this extra penalty to

to the training process um to penalize

large weights large coefficients

um is known as regularization.

Okay. Um and there's two common

penalties that are added. Um so there's

actually two different variations on the

penalty. Um we're going to study both of

them and um they're they're known as

lasso. So if you take linear regression

and add a particular type of penalty,

it's known as lasso. If you add another

type of penalty, it's known as ridge

regression. We're going to study both of

those and what their differences are.

But these are the primary two

uh regularization tech uh models that

are used um to take a regular both of

these take regular linear regression and

just modify the training process a

little bit in different ways. Two

different ways. um using that alpha

um to penalize the terms in slightly

different mathematical ways. So we're

going to learn about these two guys.

Lasso regression there. Both of these

are just offshoots of linear regression.

So underlying model is still linear

regression. It just adds different types

of penalties to the training process.

So both of these are still in the family

of linear regression. In fact, in um in

scikitlearn, they both come from they

both are still from the linear model

family in inside of the linear model

module, which is where linear regression

comes from. So there's still linear

regression. They just have different

styles of penalties added to them. Um

which we're going to see.

Okay, so just to recap that

regularization is the process of adding

a penalty to the training to discourage

complexity. In this case, we're going to

discourage large coefficients.

And um this should help prevent

overfitting.

And so uh these are going to lead us to

two different offshoots of linear

regression that have two different

penalties.

lasso and ridge regression, which we're

going to uh study next,

but they they function the same way as

linear regression. They will just have

different penalty terms added onto their

training process um to discourage

uh discourage um again those large

weights.

Okay, any questions about regularization

before we first look at our we're going

to look at our first uh variation on on

our first regularization technique which

is going to be called lasso regression.

Okay, let's look at lasso regression. So

what is lasso regression? It's actually

lasso is short for least absolute

shrinkage and selection operator

regression. Um and this will function by

adding a particular penalty to the

linear regression model. So again, it's

based on linear regression. That's the

underlying model. It's just that during

the training process, we are going to um

add a penalty which has the effect of

shrinkage of the weights. That's why

it's called shrinkage. It encourages

smaller weights through that penalty.

And it also will shrink some of them so

much that they'll become zero. And so it

has has an effect of kind of selection

which means that some of them get wiped

out to zero.

And this means that whatever is left

over is kind of what's selected as our

features because the other ones will

have zero weight applied to them. So

this penalty will really favor small

weights um and penalize really large

weights. In fact, it will favor small

weight so much that some of them will

actually um be shrunk to zero um during

the training process. And the ones that

are left over are the ones that um are

the ones that are what we call selected

because they are the ones that remain in

in the training um after the other ones

get uh coefficients of zero. Um now when

you make some of the coefficient zero

you are inherently making the model

simpler right there's less features

involved in the prediction that or less

features that have an effect on the

prediction. So this definitely makes the

model simpler. This lasso this shrinkage

and selection uh process makes makes the

model simpler for sure. Um

and this is supposed to reduce

overfitting. Right? If you make the

model simpler, it's not as complex. It

has less of a chance of memorizing

training data and not generalizing over

to test data. So our whole goal with uh

regularization is to make our model

better at generalization, right? Over to

test data from the original training

data.

Um so how does this happen? We have to

go back to the

uh training process. If you guys

remember, I I wrote out this equation a

little bit earlier, which is the

distance. This is the sum of squared

distance between our labels and our

prediction.

This is basically the mean squared error

uh calculation that we're trying to

reduce when we build our model using the

training data. Um so this is just in

standard linear regression. This is the

um uh sum of squares uh distance, right?

So this is this is what the model is

trying to minimize when it learns these

coefficients.

So when it learns these coefficients,

it's trying to minimize this guy

minimize. It's trying to find the betas

that minimize this quantity

mathematically. That's what it's doing.

Um and there's there's a algorithm that

will discover what the best betas are

that actually minimize uses that gives

us a line of best fit, right? That's

what we've been talking about for

regression.

Now, in regularization,

here's, by the way, here is that same

thing, but we've just inserted our model

for the predictions. This is our model.

Just a fancy way of writing down our

model, right? It's the beta 0 plus all

of these betas. So, beta 1 x1 plus beta

2 x2

plus on and on and on, right? That's

that's what this uh means. If you're

unfamiliar with the sigma notation, it

just means sum. So it's the sum of all

these guys or this term. Um, so this is

this here is just a regular linear

regression

uh training regular linear regression

training. So we the training process

solves for these parameters, right? It

solves for these weights. We discover

what those are by minimizing this

quantity. That's the whole training

process. Um, but when we do lasso,

we add a penalty which is this.

Here is our penalty.

So basically um we take our linear

regression training which is this and we

add on a penalty which is this. And you

can see exactly what this penalty when

when you minimize this penalty. It's

when these weights are small. So this

encourages

So minimizing this quantity encourages

small weights

encourages small betas

beta I

right you or in this case beta j sorry

this encourages small beta js uh because

we want this thing to be minimized

minimized

so Um, what's going to make this minimal

is of course the line of best fit and

small weights, right? Are going to make

are going to bring this error down the

most.

So, um, and here's our alpha, right?

Here's our alpha. So, you can encourage

a higher penalty with a larger alpha or

a lower penalty. If alpha equals zero,

what happens to that term? It just goes

away. So if alpha equals zero, there's

no penalty and we're back to uh we're

back to regular

linear regression.

We just have regular linear regression

because we have no penalty at that point

when alpha equals zero. So the smaller

alpha is, the less penalty we're

enforcing and in the regularization.

Okay.

Now what happens is in reality when you

train with lasso. So this is lasso is

this particular penalty. This is called

the lasso penalty

or sometimes um people call this the L1

penalty.

Um L1 just comes from the fact that this

is the first power or absolute value. Um

so it's not a squared penalty, it's a

single uh single power penalty

um there. But when you add this lasso

penalty, what can happen is it it does

because the because you're minimizing

this, it does encourage some of these

weights to become zero.

So some if you're really trying to get

the lowest quantity of this,

the lower the better.

What makes this thing lower is of course

if some of these go away if some of

these go to zero then that of course

will lower this as much as we as much as

possible right so what happens during

the training is some of these

coefficients actually they're encouraged

to be small because of this penalty but

some of them will actually become will

actually become zero um in order to get

the best model the best fit some of

these will actually get so small that

they'll basically become zero

And that means that that that feature

basically has no effect anymore. It's

it's been the model has been simplified,

right? That feature no longer really has

an effect.

So just to call out the alpha again, um

if alpha zero some code, uh basically

you have your linear regression, you're

back to linear regression because alpha

0 is just wiping this out and you're

back to linear regression.

um if alpha is infinity. Now if alpha is

infinity that's an extreme. So if alpha

is infinity the only way to make this

minimize is if all your coefficients are

zero. If every beta is zero then this

will lower the the error as as much as

possible. So you basically have no

model. So if all coefficients are zero

you have no model and that's useless. So

you don't want your penalty you don't

want your alpha to be huge is what this

is saying. You also don't want your

alpha to be small. you're basically back

to linear regression. So you want

something in between. Um and the typical

typical value is alpha equals 1.

Typical is alpha equals 1

to have some level of penalty there. So

just a regular kind of regular penalty

term.

But we are actually going to have a way

to test and evaluate which alphas are

the best.

Um,

basically you can yeah you can have a

you can have a penalty that's close to

zero. You can get rid of this if just a

regular linear regression performs

pretty well. You can basically have no

penalty in that case.

Yeah. So nearer zero or like it could be

that adding a little bit of penalty

actually helps the overfitting and it

could be really small. One thing that

we're basically going to do is have a

strategy to try out different alphas.

try different alphas

and evaluate performance

and then we can decide which so that's

what we're going to do is have a

strategy to just plug in different

alphas generate the like train the model

and then see what its performance is and

see if those alphas are good what what

which alpha is the best we can evaluate

that

because we can train the model and see

what it performance is

right.

Yeah. Yeah. So, we'll do that. We'll

practice that.

Okay. Great. Any other questions about

this lasso regression? So, remember this

is linear regression here. This is the

this is how you're training to find the

betas in linear regression. So this is

just linear regression uh um training

function there.

We're adding a penalty which is this is

the lasso penalty

lasso penalty there right we're adding

that this is known as regularization

and the goal of regularization is to

prevent overfitting. So you add a

penalty here this makes the model

simpler which prevents overfitting.

helps you generalize better when it's

simpler.

Any questions conceptually on this?

We're going to do a code example with it

coming up, but any questions on this?

Uh yeah, you you so that's the thing,

Ronald, is you may be willing to

sacrifice some accuracy in order to

generalize to unseen data because

remember that's what we're really trying

to get after is we may be willing to

sacrifice some accuracy on this training

data in order to have it perform better

on the test data, right? we may be

willing to do that. That's a willing

that's an okay sacrifice

as long like if if it generalizes

better. That's what we want. That's what

we're trying to do here is add a

penalty, make the model simpler, and

help it generalize better to new and

unseen data. Right?

That's that picture I've been using with

the with the um train and test split.

Where is the square?

So in the model there's no square. So

remember the model is the model is this

um equation uh that has no squares in

it, right? It's beta 0 plus beta 1 x1

plus beta 2 x2 plus beta n xn.

That's the that's the linear regression

model. This is the now this this is the

model but this is the equation that

helps us train and find the betas. This

is how this is what we find the betas

with. So we'll continue. Um we were

talking about the lasso regression which

uh adds it takes linear regression right

which is this optimization and adds in a

penalty um scaled by the alpha. Um, and

what that does in order to minimize this

whole thing, it encourages these to be

small uh as possible. Um, which makes

the model simpler, right? The weights

don't get overly big and complex. Um,

they they tend to stay small. In fact,

some of them can even go all the way to

zero. Um, which makes the model even

more simpler,

right? Um, so let's practice uh using it

in code. It's actually really easy to

use. It's going to be essentially the

same uh style and and code as linear

regression except we are um just going

to have to uh put in our alpha parameter

um when we use the lasso. So here we are

um from the linear model family right

which makes sense. It's a linear

regression offshoot that has this

penalty in it during the training. um we

are grabbing our lasso regression. Um it

also has a version of the lasso that

we're going to take a look at that is

used for cross validation which is

really um convenient as well. So it has

a cross validation lasso which is a

really convenient um combination of

basically cross val score and lasso um

all in one. So it actually is really

nice to use that way. Um so we'll take a

look at that example. Um, but we are

importing it. The main thing is going to

be the lasso model here. Um, we're going

to be using a different data set for

this one. So, not the ocean uh data, but

this hitters data, which is a baseball

data set. Um, so it has 322 rows um with

20 different columns and it looks like

this. So, you want to download that one.

Um, hopefully you guys have access to

that one.

Um,

so I will upload it into

this.

So give me a moment.

There's that. And then we can run this.

Okay. So we are displaying the data and

so it has um the the hitters names and

then it has a bunch of different

statistics. These are all baseball

statistics.

Um, if you're unfamiliar with with them,

that's okay. It's not a big deal. Um,

but just different baseball stats here.

Okay. Were you guys able to load that?

Um, if you're following along, were you

able to load that? You should have

access to this data. The hitters CSV.

This is the one we're going to use for

the lasso model

to build a lasso model.

Yeah.

Okay. Able to load that one. Perfect.

Okay. So, able to load that one. Um, and

we take a look at the the head. Um, so

we're actually going to uh drop this

unnamed column because we don't care

about their name. it's actually just the

batter's name which is not going to be

useful in modeling. Um so and remember

that's generally true like an ID, a user

ID, like a customer ID, a name, that's

usually not going to be useful in any

kind of modeling. So we're actually just

going to drop that uh column and we're

going to do it in place.

And access equals 1 means we're dropping

that column. Um, so we're going to drop

that and we should no longer have that

column and we have all of these guys

now. So you want to run that. This will

drop that. Um, this will drop drops the

column in place.

Um, and now we can see we have uh all we

have this data where um we have this

data where it's now removed. So, this

that column is now gone and now we have

these guys. Um, do you notice anything

about this

from the info?

Looks like we have a couple categorical

features, a few of them, league and

division

and new league. What do you notice about

this

nullles? Yep. So, there's definitely

some missing data there um that we're

going to have to deal with.

So, it looks like there are uh there are

59.

Um there are 59. Now we could we the

alternative to doing that is we could uh

we could just use our usual code where

we do dfis

uh isnull.

Um and then we do uh dot sum to total

those up across our different columns.

And we can see that uh we have 59 of

those in the salary column. That's this

is the standard way of doing that,

right?

standard way of doing that. And we have

so we have 59 of those.

59 of those. So we have to deal with it.

Any ideas on how to deal with it?

Any ideas on how to deal with it? This

is now this is 59 out of 300.

So,

what do you guys think about that? It's

a little bit different than 200 out of

20,000. A little bit different. We have

We have about 60 out of 300.

There's a decent amount.

Any ideas on how to handle this one?

Replace. Yep, we should replace. What do

you think we should replace with?

It's a float. It's a floating point uh

value.

By the way, something unique about this

that's a little different than usual,

too, is that the uh this is actually the

column we're going to use as our label.

So, we're actually going to predict the

salary based on the uh based on the um

rest of the features. So, we definitely

need to fill in these nles, right?

Because they're actually going to be the

labels.

We're missing some labels uh in our

data.

We definitely need to fill them in.

Yeah. So, we're going to replace them.

All right. So, we'll we will replace

them down below. That's going to be

coming up. Uh we'll come back and

replace them. um before we replace them,

we're actually going to get our uh one

hot encodings for those three different

um features we have. Um so we do uh get

dummies with this. Now um of course we

don't need to do this if we just so this

code we don't need to do if we just pass

in the dype here

um which is uh then we don't need to do

this. So we can comment this out.

Um so now what I want you guys to notice

is this is the alternative to what we

did before where we are purposely just

doing these columns not the whole data

frame but just doing these columns and

then we can um concatenate those these

one hot encodings. We're going to

concatenate back to the data frame.

Right? So if we do our dummies and then

do dummies.info info. Um, we can see

that we end up with six new columns. And

in fact, we can do dummies.head

and take a look at what those are.

Right? So, these are league A, league

uh, N, division E, W, division W, new

league A, new league N.

Okay.

So, um these are uh these are our one

hot encodings for these three different

features which are strings, right? So,

those those features were strings. If

you go back up, those were our only

string features we had. So, we've one

hot encoded those so we can use them in

our model. What we need to do is just

concatenate this back to our data frame.

Right? So, we just need to concatenate

it back into our data.

Okay. So, what we're going to do then is

we're going to grab um we're going to

grab Y as our salary. And of course,

we're going to fill nles on that Y

coming up shortly. But we're going to

grab Y as our salary and X new. Now

before building a full X, we're going to

take a look at X numerical as our data

frame minus these columns. The reason

we're doing minus those is because we

are going to concatenate our dummy

variables back into this that are going

to replace these guys. So we're going to

replace these anyways with our one hot

encodings. We don't want the strings. So

we're going to get rid of those. And

we're also going to get rid of the

salary because that's going to be part

of our that's just a label. So we don't

want that in the X, the eventual X.

Are you guys able to run this one?

Hope I'm not going too fast. You guys

able to run this? And does it make

sense? What we're doing is we're putting

our labels in Y, which is what we

usually do. So, we're going to predict

the salary

and we're getting ready to build the X.

But before we first want to get rid of

those one hot the strings. This is

getting rid of the strings

and this is getting rid of the label.

And that's going to be part of our

features. What we need to do is build

our final X by concatenating our dummies

with this. Do you guys see that? We're

going to concatenate our dummies with

this to build our final X.

But but prior to doing that, we need to

get rid of these string columns here. So

we're dropping those

dropping those from the uh data frame uh

and getting a numerical uh x numerical

here.

You can see the columns of that are just

these guys here. So the the results we

need to concatenate our we need to

concatenate this guy um into this and

then that'll be our full x all of our

features.

Okay. So you can see x is going to be

pd.con

of this with our dummies.

This with our dummies. And um

uh instead of doing this, I'm actually

going to do the full dummies. We don't

need to

um pick just a few columns. We're

actually going to do our full dummies

here and um do x equals 1. Now, the

reason that's the case is because um

this will get rid of one column per

feature and basically assume that if you

have a if you have a zero, the other one

should be a one. If you have a one, the

other one should be a zero. Um so it

basically makes that assumption because

we only have two of them. Um so whenever

there's a one, the other should be zero.

Um, so you can get away with just having

these three, but um I think it makes

more sense to just have to have the full

dummies,

but by process of elimination, you can

get away with just using two of them

because anytime you have a zero, the

other one should be the other feature

would have been would have been a one,

right? And vice versa, when there's a

one, the other feature would have been a

zero.

So we do that one.

And you can see all of our uh all of our

one hot encoding features end up back in

there.

So this is the code that I want you guys

to run. I think it makes more sense. It

follows along what we've been doing.

um which will concatenate our dummies

back to our features here to build out

our full X. So now X is all of our

features. Um remember X

X contains all of our features

now.

So X contains all of our features and so

we have all of this now.

Okay. Were you guys able to run this

one?

Damn. We have y, we have x. We still

need to deal with the nles in y. So that

something we still need to deal with.

But hopefully you have this. Now

all these are numerical.

So that should be good with the model.

That's one thing about X is you should

you our X should have all numerical

features, right? Because it's going to

go into a model to to learn those betas.

So it needs to have all numerical

features,

right? These are going to be all

numerical, which makes sense. We change

we did one hot encoding to change all

those guys to numerical.

Sorry, I'm scrolling down.

Okay, we do fill in the nil later. Okay.

Okay.

Any questions so far? So, we're just

getting our data ready. We haven't

applied the lasso yet, but we're just

doing some prep. Now, hopefully you guys

recognize th these are some standard

steps that we're taking when we do our

modeling. We have to do these data prep

steps. They're necessary. And so, if it

seems like it's a lot of work, that's

because it is. It is work that you do to

prepare your data to get ready for

modeling. You have to do that. Okay.

So, we're doing that here. Um, now we're

going to do our train test split because

we're just going to do uh we're going to

do hold out here. So, we're doing a

train test split with about with a test

size of about 0.25. So, again, anywhere

between 0.2 to.3 would be okay.

Um, so uh

it's our choice. We could do 0 2. We

could do 3. We could do anywhere in

between there. We're doing 0.25. That's

fine. Um, that's okay. So, we we build

our train test split right there. Um, so

pretty pretty simple and we've seen that

a bunch of times with our X and our Y

data frames. There we have our train

test split.

Okay.

Um, now what we're going to do is do our

our scaling. So, we're we didn't do this

last time, but we're going to do this

now as uh because we should get in the

habit of doing that. Um is um we're

going to um go ahead and scale our

features and we're going to use the

standard scaler here uh to do that

scaling. Okay. Now, we could use minmax

scaler. That's fine, too. We're just

going to use the standard scaler here.

Um and remember we are going to uh um

use the standard scaler from sklearn and

we're going to transform our features uh

uh according to our um according to our

training data. So we have our

pre-processing standard scaler here. So

we import that guy and then we um build

our standard scaler and fit it on the

training data only on the numerical

features. Um so that's which is going to

be uh all of these guys. So we're doing

the scaling on all of these guys. Now

something to note is that we are not

scaling all of these one hot encodings

mainly because it doesn't make sense to

scale those really. They're zero or one.

They don't need to be scaled, right?

They're already zero and one. So they're

they don't need to be even if we were

doing minmax scaling, it's going to put

them between zero and one. It wouldn't

affect it really, right? So these one

hot encoding features, we're not going

to scale because they're they're always

going to be zero or one.

There's no need to scale them really.

Um, but we're going to scale all the

other features here that are floats.

So that's these guys here. These

numerical features we're going to scale.

Okay.

Don't need to we don't really need to

scale the one hot encoding. Uh, it's

pretty much already scaled.

Oh, you should change that. Um, go back

and rerun go back and rerun this, but

make sure you have your data type as int

here.

Make sure you add that in there to

change that over to integer and rerun

that and then rerun the rerun the

concatenation.

So, make sure you run this

and then u make sure you rerun this and

rerun the concatenation part which is uh

this

Okay. So, we go ahead and fit the um

scaler to this data and then we're going

to transform our training features,

those numerical features. um we're and

then we're going to uh transform these

features uh uh the test features in the

same way. So we're going to perform the

same transformation from the scaler on

the test data. So that's something

really important I want to note here is

that we always scale both the training

and test data. We always scale both. Of

course, we're going to train the model

on the training data. Um, but we are

going to also test it on the testing

data and it also needs to be scaled

because our model that we build is going

to assume scaled features. The

coefficients that it learns are going to

be assuming scaled features.

So, we need to also scale our test data

in the same way. So, we're doing that as

well.

So, we scale that and now we have our uh

training and testing features have been

scaled.

No, we haven't replaced. We're going to

do that. We have not yet. We're going to

do that coming up in a minute. Yeah, we

haven't done that. Um, it is it is the

label. We definitely need to replace

NLES. We just haven't done it yet

because it's not in the features and

we're doing all of our uh uh

pre-processing to our pre-processing to

our features.

Yeah. So, we're definitely we need to

we're going to in a minute.

Okay. So, if you look at the data now,

it's all been scaled. So, these are all

um zcores. These are all on a much

better scale now. Um, and these are we

still have our one hot encoding features

which are zero or one. So this scaling

should lead to a better model than if we

didn't scale. So scaling is really

important. We can see that here.

Okay.

Now, um, let me ask you guys, were you

able to run the scaling? Are you caught

up to here? If you're following along,

were you able to run the scaling?

Okay, great. Great.

Awesome.

Okay. So, uh what we're going to do now

is replace NLES in the uh replace NLES

by calculating the median of the data.

So, what I want you to notice is that we

are taking the NLES now this is um this

is on purpose is we are purposely taking

the NLES um out of the median

calculation. So we're skipping the NLES

when we compute the median because we

don't want those NLES to affect the

median calculation.

Um so we compute a median salary here

and then we fill our NLES with the

median salary um from the training data.

So this is our choice. This is a choice

um to use the median and it's also a

choice to use the training set median

for both train and test. What we could

have done this is an alternative that we

could have done is use the entire column

and then um use the median of all of the

data to replace. That's really up to us.

Um this is one way of doing it. We could

have done before we did the split. We

could have um filled in with the median

earlier. We chose to do it here mainly

because it doesn't affect the features.

So we could have done this earlier and

did it before we did the split and

filled the NAS. Um really doesn't it's

doesn't matter that much which way we do

it. Um but we do need to fill in NLES.

We cannot have those be null when we

when we put it into our model. So some

way we need to fill in nulls. Um and so

in this strategy we're filling in our y

train um with the median salary from our

training data. And same with this we're

filling in with the median salary of the

training data as well. But that's a

choice. We could fill in with the mean

with the average. Um we could fill in

with the we could do it with all the

data together before we split it. we

could have filled in with all of the the

median across the whole data set. Um

either one works. You can do it either

way, but we we did it um later here to

show that it doesn't really affect the

features. So, we can choose when we do

it, right? It doesn't affect the

features at all. So, we can do all of

our pre-processing on the features and

then do our label uh filling in NLES um

if if we have them.

uh x numerical. Um make sure you're

running uh this

uh x numerical was defined here

when we split it apart um from

uh when we dropped these columns here.

So make sure you're running this. This

is x numerical.

It's defined there.

So go back up this uh this cell

where we split apart the y and we and we

have the x here x numerical.

Make sure you run this.

Make sure you run this. And then you can

run these. Then you run this to build x.

All right.

Are we up to here with this filling in

the labels?

Uh because then we can build our model

once we're up to here. We've scaled

everything. We filled in our NLES.

We've gotten one hot encoding.

Yeah, it is. That's why you know that's

why we spend a lot of uh time on model

on data preparation with pandas, right?

That's why we did all that pandas work

for sure. Yes, there is a lot of work

before we can build a model.

Yes, the mo do you guys notice that like

the modeling is relatively easy. It's

just a fit and predict. The modeling is

actually really easy. It's all the other

work that's that's more involved, right?

more code.

The modeling itself is really easy.

It's just it's just one line of like

ffit.

Yeah, pretty easy to do.

And then you do evaluation which is a

couple lines.

Yep. There's these are all the these are

the common steps. All these steps we're

doing are very very prototypical in

model building is you let's just go back

through this to see what we did, right?

We imported our data. Um we analy we

dropped this name column because it's

not useful to us. So we dropped that. Um

we filled in the NLES eventually. Um but

you know if there were any nulls in our

features we would have to deal with

those as well by replacing them or

dropping the rows like we did earlier.

Um

and then we do one hot encoding because

of course we can't have any string

columns going into our models. We got a

one hot encode.

Um we uh then build our X and Y by

concatenating the one hot encoded back

to the numerical features.

Then we train test split. Right? That's

pretty common. Or we could do cross

validation either way. Um a kfold cross

validation. Then we scale. So we didn't

do this last time, but this is something

we should get in the habit of is scaling

um our features. So we do that. And now

we're ready to model. So now we're ready

to model. Um so that's this part.

Okay. So let's do the model. Um the

model is actually uh pretty easy to do.

So we're going to use a lasso. So we

have a lasso model here. Notice what

we're setting our alpha to. So the big

parameter we really need, ignore this

iterations. We actually don't really

need the we don't really need that

parameter. Um so just ignore it for the

moment. But the big one that we're

setting here is the alpha. So when we

did linear regression, we didn't need

any parameters to go inside the linear

regression object. We didn't need any

parameters, right? Because there are

really no parameters of it. But for

lasso, the important one is the alpha.

And so we need to know what to set alpha

to. Um let's start with alpha equals 1.

That's a good starting place. So a

typical um starting point

for alpha

um

is uh is one. So that's a typical

starting point. And so we can set alpha

equals to one. This max iterations is

the the parameter that governs the

training process because it is

iterative. So if for some reason we we

can't converge to the right betas and

we've run it for 10,000 steps once we

pass 10,000 steps, uh it will stop and

just give us the betas at that point.

But it will likely never hit this

number. It'll converge before then. So

um we don't really need to um specify

it. So, I'm actually just going to get

rid of it. Um, it's not really a big

deal. It should converge before then.

Um, but if if we want to set like a

maximum step size in the optimization,

we definitely could there. Uh, but not

concerned about that too much. But

here's our lasso. And then we're just

going to do a fit on our data. So, look

how easy that is. Just like a linear

regression, lasso.fit,

right? So, we do fit. Um,

oh, I didn't run this. I'm sorry. I got

to run this. Okay. Actually, that's a

good example of what happens when you

don't when you have nles, right? So, it

says our our null contains nan. That's

because I didn't run this. But now that

should be filled in. Now, we should be

able to run this. Okay, perfect. So it

runs.

Okay. So you can see what the intercept

is. Um this is one of our coefficients,

right? The intercept is 457. And look

now what's really interesting about the

coefficients is look at what some of the

coefficients are. Some of them are

actually zero, which is really So some

of them ended up being at zero, which is

very very interesting. that means that

those features get cancelceled out and

they're basically not part of the model

which is really interesting. Um, so we

have all these coefficients and some of

them are zero.

Yeah, negative0 is just because of the

convergence like they started out

negative and worked their way up to

zero. it. Negative Z really just means

zero, but they just were coming from

they were like small negatives and ended

up at zero

during the training process. They were

negative at one point and it ended up

zero. Um

so yeah, negative 0 just obviously means

zero. Um it's still still zero there.

So what's interesting is some of these

features ended up uh being zero which

you don't usually see in a linear

regression. So if we were to train this

using a linear regression we typically

wouldn't see that but some of these

turned out to be zero because again

we're encouraging those betas to be

small. we're encouraging them to be uh

small and so um you know what happens is

some of them can be shrunk all the way

down to zero meaning those features

don't contribute that's a really simple

model at that point right so we've taken

something complex that includes all of

these features and actually reduced it

into something simple that only includes

these features

right

so that's what it does um now we need to

evaluate this to see how good of a model

it is. But that's what this is saying

here in this text is that um a positive

uh coefficient indicates that as the

independent variable increases the

dependent variable also increases.

Negative coefficient means as the

independent variable increases dependent

decreases because it's reducing the

value. Um and lasso is known for feature

selection by shrinking some of them to

zero effectively removing those

variables from the model from the

equation right

um

so that's what happens

some of them end up being zero

were you guys able to run this this

lasso uh fit which is the training of

the lasso Control.

No, it doesn't ensure there's no

overfit, but it helps with overfitting.

It's supposed to help by making the

model simpler. And this is definitely a

simpler model because it's removing some

of the features from the model

essentially, right? Because some of the

features aren't going to contribute.

It's a simpler model.

It doesn't it doesn't mean there's not

going to be any overfitting, but it

helps prevent it. That's what it's

designed to do to help prevent it.

Yeah. So higher coefficient. Yes. The

higher coefficient means it's a more

important feature towards the

prediction. Yes. That's what it means

for sure. The higher the magnitude, the

more of a contributor towards that

prediction. Uh it is. Yes.

And it's not just it's it could be

higher positive or negative there. Like

a higher negative is also a pretty big

factor,

right? So So you want to think about it

in terms of absolute value.

does not guarantee but helps. Yes, it

doesn't guarantee it but it's designed

to help overfitting, help prevent it.

Yes, absolutely.

Okay.

So let's do some evaluation. Um so let's

do in this case we are going to do our

predict

Oh, yeah. I'm not sure why that's the

case.

Interesting.

We could try increasing the um max

iterations.

Okay, that's why. Yeah. So then you get

that result with the with the higher max

iterations.

It doesn't get cut off there.

I think that's why you probably left

this in there,

which is fine. You get about the same

numbers.

Yeah.

All right. Let's evaluate this. So,

we're going to to to do evaluation. I

want you guys to see again. We should

get in the habit of doing evaluation,

which is taking our model and predicting

on the training and predicting on the

test sets, right? So we predict on the

train set and calculate our MSE

and we um calculate our R2 score um or R

squar score I should say. Uh but again

the MSE is the one we're really going to

use mostly. Um but we calculate so we do

our predictions and then we compare that

into our mean squared error with our

labels

and we uh go ahead and do the same thing

with the test. Right? So we do uh

lasso.predict

on our test features and we go ahead and

compare that with the test labels. And

so what we're doing there is generating

our MSE.

So, we we take a look at our MSE and we

get uh 84,000

MSE. Um, and so, of course, we could

take the um what we could do with that

is take a look at the um MSE on the uh

we could do um MP. Square root

and do the square root of the MSE test.

and we get um 340. So this would be in

the units of our label. So, we go back

and look at our label um for some of

those um

so uh we are in 300s and our data is

like right around the 500. So, of

course, if we describe this um we could

see what the statistics are of it. So,

we could do df.describe describe and

generate that. But that doesn't look

like a very good error, right? If these

are in the 400s, um that's that's not a

very good error.

So again, it's not a very great model.

But one thing I want you to see is that

it's it's not overfitting.

Um if anything, it's actually

underfitting, which is what this kind of

um MSE suggests, right? because our

error here is 84 uh excuse me 84,000.

Um

our our area here is 84,000

excuse me and on the test set it's

116,000.

Um so these two errors are both bad. So

it's not overfitting. This is actually

underfitting. So it's not overfitting,

it's actually underfitting. Um, and so

that's the risk with something like

lasso is that it's making the model a

bit too simple and we actually risk

underfitting, which is what happens. We

have too much error across both the

training and the test set. Overfitting

is when we do we have really good

performance on the training set, but bad

performance on the test set. We're not

overfitting.

um we are uh underfitting because our

performance is not good either way. Even

this R squar is pretty low. It's not

even at 50%.

Okay, so that's so we we do the

evaluation and again the evaluation just

comes down to making predictions and

computing our error amongst those

predictions to our labels. That's always

what the uh evaluation is going to be

for MSE.

What's the ideal MSE? What do you think

it should be? What is So, think about it

like this. The MSE represents the

average distance between our predictions

and the labels.

So, if we're getting it right all the

time, what's that distance going to be

if we're always right? What's our

distance from what's our distance from

our predictions to our labels going to

be if we're always getting it right?

Zero. Yeah, there's not going to be any

distance. It's going to be right. It's

going to be perfectly aligned, right?

There's going to be no distance there.

So, yeah, an ideal MSE is zero. That's

an ideal MSE.

So, anything close to like the smaller

the better for MSE. The smaller the

better. Um, for this R squared, uh, it's

it's a scale between 0 to one where one

is the best. So, one would be perfectly

aligned predictions. Um, so, and again,

this this is we actually multiply by 100

to get uh because it's it's a number

between 0 and one. So we get about 47%

which is not good.

Okay.

All right. Any questions on this

evaluation?

All right. I want to show you something

which is

Yeah, this that's true. The scale of it

matters on the data because we should be

you should always interpret your MSE in

the scale of

um your your labels because your labels

like in this case our labels um you know

we could take uh for example we could

easily let's actually do that let's take

the average

let's take the average of our labels on

the training data

and and we could see what those are. Um,

so the average is 500,

right? The average is 500. And look at

what our uh square root of our MSE is,

which is in the same units as our

original. Um, so we have uh quite a bit

of error. 340 when our units are right

around 500.

So that's quite a bit of error.

Yeah, MSE of zero means our our uh our

predictions are nearly identical to the

test labels. Yes, that's what MSE of

zero means. There's zero distance.

So closer to zero, the better.

But we talked about it as you you really

so the rule of thumb should be what is

your RMSSE as a percentage of your

typical value. So your typical value is

in the 500s. Our our RMSSE is 340.

That's just really high. That's over

like 60% of that value.

So that's just a lot. That's too much

error. What we would love this RMSSE to

be is under 20% of the typical value. So

that means on average we are 20% or less

off in our prediction. That would be

good. That would be pretty good. That

means we're like 80% accurate,

right? That'd be pretty ideal. So you

got to think about it in terms of this

RMSSE, which is in the same units as

your labels.

This is the

RMSSE

which is in the same units as the

labels.

So and then to interpret this we have

340

is compared to

typical

um salary unit of 500

right so this is uh quite a bit when the

typical value is 500 and we are off on

average by 340 units

that's so much relative to the typical

value

that's just too. That's a lot of error.

That's not a very good model, right?

It's underfitting. It's definitely

underfitting.

Yeah. So, that's a great question. What

should we do from here? So, um because

we're underfitting

um we should use a more complex model.

So uh we're going to learn about those

in lesson four, but we should use

something different. This linear

regression is still too basic. Even with

lasso, it's still too basic.

Yeah, we're underfitting because we But

it could also be we're underfitting with

a regular linear regression. We should

test that out. Um, and maybe it would be

an exercise for you guys um to test that

out yourself. It shouldn't be hard to

do. Um, you already have all the data

scaled. You So, do you see how you would

do that? You would just come in here and

build a linear regression rather than a

lasso and dofit and then you would

evaluate it the same way with a

dotpredict. It's really easy to do that.

And then we can compare that um to to

this. It shouldn't be that hard to do

that, right?

And something you guys could do for

sure. Um,

is build the linear regression and

actually compare it and see what kind of

difference it makes. I mean, we honestly

we could do it ourselves. We could do it

right now. Maybe it's worth trying that.

So, let's build a linear regression

for comparison.

So we have our linear regression

uh is linear regression and then we do

ffit linear regression.fit fit

right so so this will train it um and

then we can evaluate it so lin MSE is

mean squared error

and then we can do our um let's do our

training let's do the training and then

um let's predict

actually let me do that here

uh y prediction

train

linear

equals um linear regression.predict

and then we're going to predict on our

training features.

Okay, do you guys see what I'm doing?

I'm building a linear regression for

comparison.

I'm doing dofit here to train it and

then I'm making some predictions on the

training set and we're going to evaluate

those. I'm going to replace that here

with y prred

uh train

linear. So these predictions

Okay. So, if you guys want this code, I

can paste it in.

So, let's see what the RMSSE for just a

linear model is.

It's a little bit better. It's better

for sure.

So 289 is better than this 340. It's

better. It's getting closer to zero.

It's still underfitting though,

right? And that's just on the training

set. Let's look at the Let's do the same

thing, but on

Let's change this. Let's swap this out

for um test.

And then let's do test.

And then let's do test

test.

and then

test.

Okay, so this is producing test

predictions on the test set.

We are generating an MSE test

and then we're doing MSE test

which is using the test labels and our

test predictions

and then we take the square root of that

for RMSSE and then we're going to

generate that. So it's still under fit.

I mean this is still high. This is still

high um on the test set and versus on

the training set. So it's still pretty

high. Um, even the basic linear

regression is under is still

underfitting. Still underfitting, right?

Even without the lasso, which is lasso

is supposed to help with overfitting.

It's definitely not overfitting. Um,

it's definitely underfitting,

but this is a signal that it's kind of

overfitting because this is performing

better on the training data and then it

gets worse on the test data.

Definitely gets worse, right?

Did you guys follow?

I'm just running this above I'm running

this above this. It doesn't matter where

you put it. We could uh we could move it

down.

We could move it down to I just ran I

just picked a new cell right here.

and ran it. But we could move it.

Actually, let's do that. Let's move it

down

to

after the lasso evaluation.

Okay. So, I just moved it there.

And then let's move

this down.

So, I just put it here after the um

after this. So this is the um this is

basically the objective function right

of the training process. So during the

algorithm that runs when we call ffit in

scikitlearn it's going to find these

betas right it's actually going to learn

what these best betas are for our model.

Um this is our model here right it's the

combination of betas times our features

um plus an intercept beta. Uh so that's

our model but um we penalize those large

uh weights in absolute value by um

adding a penalty term like this um where

alpha is some level of penalty that we

want to provide. Usually alpha equals 1

is okay. But um actually what we're

going to learn uh to finish out this

section is there's going to be a

systematic way we can test out different

alphas um that represent the level of

penalty we want to uh apply to lasso or

even ridge

uh regression. So that was the lasso and

um if you guys remember using it was

super easy. Uh we worked through this

problem with this um baseball data um

and we had uh

let's see scrolling down we um split out

our numerical data and we did uh we one

hot encoded our our categorical data

combined it back together. Hopefully

that um rings a bell there. Um and we

actually scaled our data which is pretty

standard to do is we do some type of

scaling to our features especially our

numerical features right want to scale

those in some way whether it's minmax

scale or standard scaler um want to do

that and so we did that for this example

and then we um ran the lasso regression

which is pretty easy to use. You just

use the lasso object and you pick an

alpha here. Um, again, we are going to

have a way to test out different alphas

that could be candidates and we can see

which one's the best. Um, so I'm going

to show us that today coming up shortly.

But that was that was the lasso. If you

guys remember, we did that. Um, this it

we compared that to a basic linear

regression which is just this pretty

straightforward just a fit and then

predict and then we can generate mean

squared error. Um, still not a very good

mean squared error on this data, it's

still fairly large. Um, so it's still

not, no matter which model we use, it's

still not very good, but at least we can

practice doing that comparison. That's

what we did last time. We did this on

Wednesday.

Um

and then

we saw that the effect of different

alphas we had a lasso um

we had a lasso uh cross validation

example here. So beyond just using a

regular lasso model that um scikitlearn

has a lasso cv which allows you to try

out different alphas uh with cross

validation and um figure out what the

best alpha is. Um, now we're actually

going to have a different strategy

that'll instead of just picking random

ones, we can actually um supply multiple

parameters that we may want to test um

as many as the models may support. And

in some more complex models will have

more than one parameter like lasso only

has the alpha. Um, technically it also

has its max iterations, but really the

only one that matters is this alpha.

Other models have many more

hyperparameters that we can um uh change

and so we want a way to systematically

test out those different combinations

and to see which one leads to the best

uh version of that model. Let's say the

best results. So um we're going to

explore that coming up. So we had lasso.

Um now this is where we ended last time.

We had ridge regression. If you guys

remember, this one is just a slightly

different penalty. Um,

it takes the it I drew it out for us. It

takes the same penalty we had before.

So, it has that um residual sum of

squares error, which is the main one we

used for linear regression, but it has a

penalty with an alpha and then it has

the sum of the beta squares

beta squares. So it penalizes it has a

penalty but it penalizes slightly

differently where it uses the square not

the absolute value. That's the ridge

regression. And this has the similar

effect of you don't in order to minimize

this right because our goal in training

a model was to minimize this thing

minimize this um quantity and find the

best betas that minimize this. Um so

generally yes you want to encourage

lower values but the um once you get

values that are a fraction if you square

them they actually get smaller. Um so uh

it's it's not um it's not necessary to

shrink them all the way to zero. They

will get smaller as soon as they're kind

of below one. Um so they don't encourage

it to completely go away uh like the

absolute value does. It's just slightly

different minimization. Um so what we

see with the ridge is we don't see the

features kind of get wiped out

completely like we do with a lasso. In

lasso they get encouraged to be um to

become zero because that's kind of the

only way to minimize an absolute value.

But with squares they can keep getting

smaller and smaller and smaller um

fractions and they don't have to become

zero. It's not as harsh of a of a

penalty.

Um,

so, uh, the ridge was easy to use as

well. Um, and it also has an alpha that

we can set. So, it's literally the same

exact code, just a different model.

There's slightly different penalty and

it results in different coefficients.

You notice that none of them are exactly

zero. Like with the lasso, you can get

ones that are exactly zero. We don't see

that with the ridge. You remember that?

Um so we we s pointed out that last

time. Notice the coefficients aren't

zero. Um and then we can evaluate it. So

we did our MSE calculation which is a

pretty standard thing where we use our

model to predict on a training set,

predict on a test set, evaluate those um

by computing the metric like the mean

squed error and we can see if we're

overfitting, underfitting. This is

definitely the same kind of story we've

seen with all these models is

underfitting because the error is so big

across both sets

across training and tests. So it's it's

definitely underfitting.

Um

and same thing as lasso, it has a cross

validation uh variation on it that

allows you to try out different alphas

and um do different folds. So 10fold,

fivefolds, whatever and compute the um

try to find the best alpha that way.

Okay.

All right.

Any questions on this so far from last

time from reviewing that a little bit?

Hopefully that uh hopefully that is

jogging your memory a little bit on

ridge and lasso. Um, you know, where

we're going to pick it up today is to

finish out this lesson with one more

model,

which is going to be a combination of

ridge and lasso. So, you can actually

combine them together

um in a linear fashion, those penalties.

So, you can actually have both

penalties, the absolute value and the

square. And when you have both penalties

um that's a special model called the

elastic net uh regression or elastic net

model. Um so this is a combination of

lasso and ridge together. So you have

lasso, you have ridge and then you have

elastic net which combines both of those

penalties. Um let me show you the

equation.

So here is the uh so here is the the

model. This is the same that we've

always had. This is our usual u model

fitting for linear. This is a basic

linear regression um loss function or

objective function that we're trying to

minimize to find the betas. Notice how

we have both of our penalties though

this time. So instead of just having one

of the penalties, we actually have both.

So we have the lasso penalty

and then we have the ridge penalty here.

So we actually use both of them and um

try to find a balance of minimizing

those two uh those two penalties.

Okay. And notice how they instead of

just a single alpha, we kind of have a

balance on both of them.

So we can actually weight the lasso one

more. We can weight the ridge one more.

We can weight them the same. Uh we can

um change that around as much as we

want. So they have two different weights

there um that they could be.

Um now what happens in reality is uh

we're going to see this in the model is

that um usually what happens is these

get combined into a fraction. So there's

usually a ratio of lambda 1 to lambda 2

and this is known as the um this is

sometimes known as the L1 ratio

and this is a this is a a parameter

inside the model that we'll be able to

set um along with alpha. So we'll be

able to set an alpha and then this

ratio. Um the idea is is that um the

ratio will uh allow us to control which

one is more dominant. So if this number

is bigger the um this lasso penalty will

will be weighted more. If this ratio is

smaller if it's less than one for

example that means that the um ridge

regression is more uh dominant. Um but

the so we'll have this we'll have really

this and this at our disposal and alpha

is um

alpha is kind of like a a you can think

of it as a scale that is um so lambda 1

kind of like lambda 1 plus lambda 2 um

combined to equal alpha.

So it's like our total level of penalty

um our total level of penalty and we can

set that equal to one. We can set it

equal to whatever we want. Um and so

these will be in this ratio and there'll

be a total level of penalty that we can

apply. So the model will actually use

these two parameters when we when we do

it. But that's how they're that's how

they're all related.

Okay. So ridge uses both penalties.

That's the only difference between lasso

or sorry elastic net uses both

penalties. Um so one thing I want you to

notice is that uh if we um if we want we

could set this L1 ratio all the way to

zero.

Um which uh if we do that um the only

way this L1 ratio could be zero would be

if lambda 1 is zero. So it would just

revert back to ridge regression. So it

complet if if this is zero this will

wipe out this term and we'll be back to

ridge if the L1 ratio is zero.

Okay.

All right. So we have a elastic net

model. Um now it's used the exact same

way as we did the other models in the

code. So we have elastic net um uh from

the scikitlearn linear model family just

exactly where we had linear regression

lasso ridge all of those came from this

linear model um elastic net also comes

from there and then the cross validation

version also comes from there um

so let's see so when we build our model

it's going to be um very very simple

easy stuff because it's the same code

that we always have um we just use the

elastic net. We set an alpha alpha

equals 1 is pretty standard um just like

it is in in the last one ridge that's

industry standard is one and then an L

L1 ratio of.5

that's pretty standard as well. What the

L1 ratio of.5 is is kind of a um

uh kind of a that means that the lambda

1 to lambda 2 ratio is 1/2. Um, so

that's that's a pretty standard uh ratio

as well. But again, we could set this

equal to one and they'd be kind of

equally weighted. Um, L1 ratio of a half

means that the uh ridge regard the the

ridge penalty is a little bit more

weighted uh in that in that situation.

Okay.

So uh once we have this model um we can

do ffit and we can run that on our

training data and we can um get we can

figure out what our parameters are like

our coefficients and our intercepts. Our

model will have that but more

importantly we can use our model to

predict right so we can predict on the

test set. Um let me go back and load our

data and actually run this.

So, we're going to be using the same

data that we did for uh lasso,

which is the I'm scrolling back up so I

can load it. It's the baseball data

here.

Um,

just run it from there.

It's this hitters.csv. So, hopefully you

have that one.

Let me load this.

Okay, so we loaded that and then that

should load.

Drop that unnamed column.

We will get our dummies

and then concatenate those split

scale. I'm just rerunning things. I'm

rerunning things so we can see our model

one more time.

So rerun that. Take a look at that. That

looks good. and then

fill in the nles on the on those.

Okay. So, we should be able to run our

uh elastic net now.

Okay. So, let's import that and then

let's build our model. So, there we go.

We build our model and the intercept is

that. Now, of course, we can look at our

coefficients. Let's look at that.

Look at our coefficients. So remember

the coefficients are the uh betas. These

are our betas that are in our model. Um

so we can take a look at those. Now um

they're it's somewhere in between. It's

not a full lasso where we're going to

see some of these be zero. It's not a

full ridge. Um so the coefficients we

get are different. They're somewhere in

between there those two models that

we've already built. So not quite the

same um somewhere in between there.

Um and then we can use our model to make

predictions and and compute the MSE

uh or the RMSSE I should say as well. So

we can take the mean squared error, pass

that into the square root and comput the

RMSSE. So still pretty bad. Um this is

right around that 300 range of what

we've gotten for our other RMSSE. So,

it's not like elastic net is any better

than those other like linear or lasso or

ridge. And that's not surprising because

it's just adding those extra penalties.

We don't expect it to magically get

better. It's actually a more complex

um when we add when we add those in,

we're actually reducing it and making it

simpler. And we need something more

complex, I should say. So, we're making

it simpler um by by making penalizing

our weights a little bit more. And so

it's still not a good fit. That's not

really surprising, right? It's still not

really a great fit.

And we can we can even double check

that. We know our RMSSE is pretty bad.

Um but we can double check it with this

R2 score. And it's, you know, still not

good. Remember, a one would be really

good. Um that'd be like a perfect linear

model. This is um still pretty bad.

Okay, so as we said, the alpha controls

the overall strength. Um, so the higher

the alpha, the more overall penalty

we're supplying, which makes the model

simpler. Um,

uh, but the L1 um ratio determines the

mix or that ratio of the lambdas, the

lasso to the ridge. Um, if you have it

be um exactly zero, you you revert all

the way back to um if you if you put it

at zero, you revert all the way back to

ridge one would be all the way to pure

lassos. Somewhere in between like one

half is is good.

Okay,

so this is another example of trying out

different values of alpha in the CV to

see which one works. Now again, I'm

going to show us in a minute a

systematic way to do this, but this is

just trying out um different alphas that

we set up in this uh in this um

uh range. So we have different uh values

between minus2 and two um

logarithmically.

Um so these are uh logarithm values that

are between this between minus2 and two

and we choose 100 different alphas and

then we choose 100 different um L1

ratios between 0.01 and one and we run

that we run this um cross validation

with 10 folds. So this is quite a bit.

So we're doing 10 folds and we're trying

out 100 different um options. Uh every

time we do an option we're trying out 10

folds to evaluate it. So, it's going to

take a minute to run.

It's still running here. But again, what

this is doing is trying out different

alphas and it's it's going to do a cross

validation. And you guys remember the

t-fold cross validation is where we take

our data and we divide it into 10 folds

and then we um train on nine of those

and then test on the remaining fold and

then we rotate all the folds 10 times.

and that we average those mean squared

error metrics together um against those

10 different uh fold options to generate

a basically like an average performance

for that value of alpha. And we're doing

that 100 times for all these different

100 alphas that there are and 100

different L1 ratios that we're trying

with them.

So that's quite a bit of processing but

uh it did finish.

So we can see what our best alpha is and

our best one ratio. So we get the best

alpha is this best one ratio is this. Um

and therefore we can uh build a model

with those with just these two guys as

the alpha and the L1 and um see how that

performs.

We build that model and then we predict

on the test set and we generate the

RMSSE. It's just a little bit better.

It's still not It's just a little bit

better, but it's still not good, right?

It's still 338. It is just way too big.

Remember, this is RMSSE, so it's in the

units of our uh target variable. So,

it's in the units of, if we go back to

our data, actually, I can just print it

out here.

um this RMSSE.

If I just do this, we could take a look

at um DF

or I could look at Y test

and you can see some of these values.

These are these salary values in the

hundreds, right? Some of them are in the

thousands. Um but an error of like 338

is just too big. That's a really big

error. That means we would be off by an

average of 300 when our our values if we

just do the mean

um

is only 550 as on average is 550 but we

have this amount of error on average um

so that's just a way too big of a

proportion of error right it's not a

very good model and again we can verify

that by looking at this R2 for.

So if we go down here,

still not very good.

Here's some of our coefficients. So

remember, you can always take your

coefficients and line them up to your

your data columns. Uh so that you can

get a sense of what coefficient belongs

with what feature. So that's all we're

doing here is just creating a series

where those coefficients instead of just

printing out the coefficients, we're

actually lining them up to the columns.

So this tells us um remember the larger

it is the more influence it kind of has

on the on the final result. Um either

way, so like this has a big negative

influence. Um, this has a large positive

influence.

Okay, let me pause there. Any questions

about the

elastic net model?

This is a really this model is a really

good one to use when you are building a

linear regression and it's performing

well, but it's overfitting. This is a

really good one to use because you can

balance

lasso and ridge you can get the best of

both worlds. So the the main strategy is

if you are using a linear regression and

you see overfitting

um meaning that it's performing decently

so on the training set

it's performing okay but then on the

test set like you know it's it's not

underfitting. it's performing pretty

well on the training set, but then on

the test set it's um performance is much

worse. That's overfitting. If you're

overfitting, then this is a great model

to use because we can try basically by

by rotating through different alphas and

different L1 ratios, we can try out

different strengths of penalty and

different variations on lasso and ridge

together. This is a really good model to

to use for those overfitting cases where

linear regression is doing decently

um but it's overfitting.

Right? So far we haven't ran that case

because so far no matter what model

we've used it's always underfit. So

anytime we have those underfitting cases

it signals that we should likely just

use a more complex model and we haven't

learned about those yet.

um we will coming up in lesson four, but

um that's for this data. That's ultim

ultimately what we'd want to do is

probably use a more advanced model

because it's underfitting um just using

a linear regression and and then using

the the overfitting variations of linear

regression like lasso ridge and elastic

net.

Okay.

Any questions on this on elastic then

the TV? Yeah, we Yeah, I think I have

it. I can share it with you.

I said that and now I can't find it. I

thought I had it.

I don't have it. I thought I had it, but

I don't.

If anyone does have that one.

Yeah, I'll look one more time. I thought

I had that one.

Um,

yeah, it's not in there. I had it. Let

me see.

Yeah, I don't have it either. I thought

I had it in here.

Yeah, I don't have that one.

I don't have that one. and I'll have to

find it. Uh I have this marketing data.

I don't think this is the same one.

I have this marketing data. I don't

think that's the right one, but you can

take a look at it.

No, we're using So, for this example,

we're using the same hitters data set

that we used earlier for lasso.

No, that's an earlier one.

That's from the uh very beginning of the

notebook. So that's the that's from this

one.

Oh, this Oh, this is where it is. Sorry.

This is where it is. You can find it

here.

That's right. It was from a URL.

It was used in the very beginning of the

notebook.

And we did we did this.

Okay.

That's right. That's why I didn't have

it downloaded.

Okay.

All right. Any other questions on the

elastic net before I move I'm going to

move on to uh finding those a systematic

way to find the best hyperparameters.

Um, I'm going to show you a couple

strategies to doing that. Um, so far

we've just ran CV with some random

choices. Um, I'm going to show you a

better, more systematic approach. That's

kind of the industry standard for doing

tuning. Um, so I'm going to I'm going to

show you that next. But any questions on

the elastic net?

Okay. And again like you know

scikitlearn makes it really easy for you

guys because it just behaves the same

way as any other model. You use the

object and then you do ffit and predict

right? So the ffit is going to train it

um and the predict is going to allow you

to use that model to predict. It's it's

super easy that way. Every scikitlearn

model is like that dofitit and predict.

So it provides a really simple way to

use basically every model.

Okay,

let's talk about let's finish up this

lesson with a couple things. Um, one of

those things is going to be

hyperparameter tuning. So what is this?

The hyperparameter tuning is a

systematic way to find the best

parameters in a machine learning model.

So a lot of machine learning models have

what are called hyperparameters.

These are not the betas that we learn

during the training that's learned from

the data. These are settings that we set

ahead of time like the alpha. That's a

perfect example like alpha L1 ratio in

in the elastic net. We set those up

ahead of time and depending on what we

pick for those we get different

performance, right? And so what we

really need is a systematic way to find

the best settings for those

hyperparameters as we are training our

models. Um the the the main like idea

behind this process though is going to

be to systematically try out different

combinations as many as we want to try.

And so we're we're basically going to

have a strategy for tuning that is going

to exhaust all the combinations of those

hyperparameters that we want to try

until we find the one that performs the

best. Um and and that strategy is known

as grid search. Um and essentially what

it does is it sets up a grid um where

which is basically like a matrix to say

okay which parameters do you want to

try? I want to try um alpha and I want

to try L1 ratio

um L1 ratio like let's say I want to try

these two. So we set these up in a grid

where we say, "Okay, I want to try this

value. I want to try this value. I want

to try this value. This one, this one,

this one, and on and as many as we want

to try." So we could set up set those up

systematically like a linear um a

linearly spaced like I want to try every

alpha between between 0 and 10 spaced by

one. Um whatever, you know, we can set

up different ranges of those, but that's

going to be in this grid. And then the

L1 ratio, same thing. We can try out

different values of these that we want

to try. Let's say there's many of those.

Um maybe every um tenth between 0 to one

I want to try out. Um so you set up your

parameters and you can set up as many as

you want in the grid. And then

essentially what you're going to do to

do grid search is you're going to work

your way through every combination of

those. You're going to try out this

combo. You're going to try out this

combo. You're going to try out this

combo.

this combo. So the first value of alpha

with every possible L1 ratio, then go to

the next, try out the next value of

alpha with every L1 ratio, and on and on

and on. So we're going to try

all combos

in the grid.

We're going to try all combos and we're

going to find the lowest MSE

combination. find lowest

MSE

combo.

So whatever leads to the best model is

going to be the um parameters that are

that are deemed to be the best. And the

idea is once we have found those we know

that we can use we can go ahead and

train a model with those best alpha and

len ratio and on and on and on.

Yeah, when you get an So this goes back

to the error. Remember that for a

regression,

the error is this measurement of how far

off we are, right? So if we have a bunch

of points and we draw we fit a line

through there, the the MSE is measuring

this distance, right? So what do you

think is a good distance? Like if our

model is perfect,

what's the best distance from our

predictions to the actual points? Zero.

Yes. So the lower the better. The lower

the better. Um so for an R RMSSE, the

lower the closer to zero the better.

However, the RMSSE can be it's its units

are interpreted in the units of our

target.

So what is deemed to be good is relative

to our target. Like let's say our target

is in the thousands. Like it averages in

the thousands. If we produce an MSE of

50 or sorry an RMSSE of 50, that's

pretty good, right? Because our units

are in the thousands

and we're only on average we are off by

50 units,

right? Our distance away is about 50

units. That's pretty good. So the RMSSE

is relative to your target variable.

Does that make sense? Yeah. It depends

on the target. It depends on what you're

trying to predict.

So that's why we got RMSSE that were in

the 300s for those hitters, but the

average was the average of the target

was in the 500s. So that's a really bad

proportion of error relative to the

average target value, right? If our

RMSSE was 300,

but the target was sitting in the 500s,

that's just too much error. Way too much

error, right? That's just too big of a

value. Um, our predictions are just off

way too much in terms of that distance.

So, this would be this was a bad model.

It was underfit.

We know that from the the RMSSE. So

yeah, the RMSSE closer to zero, no

matter what is good,

zero is being perfect. Um, but it to

know what's good, you need to know what

your target is on average and then think

of this as kind of a ratio to that

average target. I think that's the best

way to think about it.

Okay, so going back to this grid idea is

so the grid is just basically laying out

all possible parameter combinations and

trying them all out by fitting and

predicting until and generating an a

metric like an MSE

until we find the one with the lowest

MSE. So find the lowest MSE combination

and that will be the best

that will be the best combo. And then if

we once we know that best combo we can

use that we can use that alpha we can

use that L1 ratio and use that model

going forward. We can we can use those

parameters in our model. So this

strategy it has a name. It's known as

grid search. So it is a hyperparameter

tuning process that tries out all

combinations.

So what's the what's the uh benefit to

this is that we get to test out a lot of

different combo combos of those

parameters like the alpha and L1. So we

can be confident what the best model is,

right? So we can pick the alpha and L1

perfectly because we're trying out a

bunch of different combinations on the

data to see which one's the best. What's

the downside?

It's expensive, right? It's an

exhaustive search. So, if you have many

different parameters and you're trying

out many different combinations, it can

get exponentially

expensive

to perform this search. Okay, so grid

search is great except for the fact that

it can be expensive if you have many

parameters with with very wide ranges

that you're searching over because that

that's a lot of combinations you have to

test, right? And especially if you have

a lot of data, that's going to be

expensive

um to do.

So uh we're going to practice doing grid

search, but that is that's the pro and

con. The pro is that we get to try out

all these combinations and see which

one's the best. The downside is it can

be expensive to do that if you have a

lot of parameters um that you want to

tune for your model um and you have very

uh many different choices that you're

trying to evaluate for those and it just

creates a really big um collection of

combinations that you have to try out,

right? Um that's the only downside to

grid search.

Now on the opposite end of the spectrum

of that is a randomized search or random

search and this will basically just um

do a sampling of those parameters from

um kind of fixed uh specified

distribution. So essentially what you do

is similarly you define your range. So

you say I want to look at alphas um

between zero or sorry between let's say

yeah 0 to 10. I want to look at a bunch

of different alphas um and I want to

look at a bunch of different L1 ratios

that are between 0 to 1

0 to one and um what we do is we say

okay I'm going to restrict only testing

20 30 40 times. I'm not going to do all

possible combinations. I'm just going to

randomly sample something in this range

and randomly sample something in this

range. And so, and I'm going to perform

that experiment a fixed number of times.

So, let's say I set the uh sampling

where I'm only going to do um 20

evaluations.

And so, 20 times we're going to pick a

combo randomly. So, I'm going to pick an

alpha and I'm going to pick an L1 ratio.

L1 ratio.

And um we are we are just going to uh

sample those randomly from this range.

Um and we're going to use those and test

those out. And then it's but otherwise

it's the same as grid search. Whatever

is the lowest MSE. Um, so whatever is

the lowest MSE is the best.

So we evaluate those, we sample, we

train the model, evaluate it. Whatever

is the lowest MSE

is the best is the best combo. Now

what's the benefit to this is it's a

much more controlled experiment in the

sense that we um aren't going to iterate

through every possible combination in

the grid. where we basically set up a

fixed number of times we're going to try

out stuff.

The risk to doing this is that you're

not

you're not exploring all combinations,

right? Because you're randomly sampling,

you may get unlucky and you may not

stumble into the best. You you can make

um samples and figure out what's the

best amongst your samples, but you may

not be covering all the combinations.

Does that make sense? The grid search is

going to try every combo. The random

search is going to randomly sample those

combos.

So, it's not going to try every single

one. It's going to try a limited number,

however many you set up. Now, if you set

that number really, really, really high,

now you're starting to approach a grid

search because now you're sampling so

many of those combos that you basically

are trying them all at that point,

right?

Um

so so that's the way the random search.

So by the way both of these use cross

validation in the sense that when you

evaluate accommodation you're actually

doing it with cross validation. So when

you do an evaluation, you're going to do

probably 10 or five folds where you

split your data and then you test it on

the rest of the folds and evaluate or

train it on the rest of the folds,

evaluate it on one of them and generate

an average MSE to get your evaluation.

So every evaluation is using cross

validation. That's why that's and

hopefully you can see why this would be

so expensive for a really big grid,

right? because you're trying out many

different combinations

and every combination is going to do a

cross validation procedure. So, it's

going to train 10 times and test against

10 different folds and average those

together. It's going to be a pretty

expensive operation

for a really big grid, right, of of

parameters.

Um, but these are the two kind of

systematic approaches we have at trying

out different hyperparameters. Remember

those those things are called

hyperparameters. These are those choices

that we have before we train our model.

Um those choices we have that affect the

performance of the model like the

alphas, the L1 ratios, those kind of

things. Um we have control over what

they're going to be. This is a

systematic approach to find out what the

best

uh value of those parameters is going to

be on our data,

right?

Okay. So, before we practice this, we're

going to practice a grid search first.

Um,

any questions?

Uh, I don't know if it has a built-in

That's a good question. by time limit. I

don't know if it has a built-in way of

doing it, but you could certainly set up

like a a a loop um to like to wrap

around, you know what I mean? Like you

could set up a loop where you check the

time. If it's if if the time elapsed as

you're doing the search, if the time

elapsed is greater than the the time

limit, then you can kind of break early.

Um so it's not hard to implement that,

but I don't know if it has that built

in. I don't think it does

because I don't think it really cares

how long every evaluation takes. It's

just going to exhaust all those

especially in a grid search.

But um yeah, I there's probably a way to

manually kind of set up a time time

loop.

So hyperparameters are um settings that

we have on the model itself and a really

good example of this is like the alpha

and L1 ratio in the in the elastic net.

So they're not things that we um learn

from the data directly like the betas in

the model like those get trained

directly by doing the um lease squares

process right um by doing that gradient

descent and all that optimization.

Um so these are not learned from that.

They're actually set ahead of time. And

so what we're saying is

the best way to understand the effects

of those is to try out different

combinations of those until we land on

the best one. Right? So hyperparameters

are those options we have in the model

like the alpha like the alpha and l1

ratio in the uh elastic net. Many models

have hyperparameters. Um we're actually

going to see that in in future models

that we study. they have options that

you can set that affect their

performance.

And so this this is just a strategy to

evaluate those different options to see

which one's the best.

Yeah. So again, hyperparameters, those

are settings on the model itself um that

affect the performance of it.

And basically we have the two two

strategies here. We can set up an

exhaustive grid and search through all

of those until we find the lowest MSE uh

option or we can randomly sample

potential options, try them out and see

which one's the lowest as well. And do

that a fixed number of times. Um

sort of like a fixed number of trials

almost. um which has a risk of not

trying out every option but but

hopefully you try out enough that you've

explored the space a bit and you get

some quality choices there but no

guarantees right no guarantees you try

everything which is what a grid search

will do it will try everything

okay now luckily per usual scikitlearn

has something to manage this process for

us in terms of grid search. Um so in

that way we will not need to manage this

process ourselves. We can just rely on

scikitlearn. And so if you're doing

hyperparameter tuning um this is going

to come from the model selection module

inside of sklearn. So we're going to

import from from skarn the model

selection module. We have our grid

search cross validation.

Okay. That's what the CV stands for.

grid search cross validation. So, this

is going to do that grid search

strategy. Um, we're going to set it up

with our dictionary essentially of

choices. So, we're going to say, hey,

here's the alphas I want to try. Here's

the L1 ratios I want to try. Um, and

here's my other settings like uh how

many folds I want to use, what my random

state is for the shuffling. So, we'll

set all that up. Um,

and then we'll just run the grid search.

And then what should come out of that is

the best options for our parameters from

the grid. And then we can use those

going forward in the we can build a

model with those best options, right? So

we're really doing some evaluation here

of what is going to be those best

alphas, those best1 ratios on our data

set, right? And the only way to really

know that is to evaluate them because

they're not things that are learned

during the training. Hopefully that

makes sense, right? They're not things

that we learn directly from training.

They're things that we have to set and

then kind of evaluate and see how they

affect things.

Okay. So, we have grid search CV. That's

going to be our primary um tool to do

the evaluations of the different

hyperparameter options.

Grid search CV. Um we're going to set up

our cross validation uh object here. Now

I want you to pay attention to this is

that um it's a slightly different

version than the kfold we had earlier.

So we've used kfold before with a

certain number of folds. This would be

10 folds and we can set a random state

for the shuffling um that happens in the

folds.

But this is actually a slight different

variation on it where it is a repeated

kfold where we do three repeated trials.

Now why would we do that? It's to be

extra extra extra careful with the

shuffling.

So this what this means is we do three

different shuffles. So we do k-fold, we

actually repeat it three times with

three different shufflings. That's all

that means. So the repeated k-fold is

actually a bit beyond the just basic

kfold. What basic kfold will do will

we'll shuffle and then do our splits

into 10 splits and then train on nine of

those. test on the other one and rotate

through all the splits.

We're actually going to do that process

three different times with three

different shuffles. So this and we're

going to average 30 results instead of

just 10. So repeated kfold is just going

ab above and beyond to do extra to

repeat the kfold three different times.

In this case only three. We could do

more.

But um now is that necessary to do? You

could argue not necessarily. Um but it

just provides extra robustness

uh beyond just our single shuffle and

then split and then rotation of those

folds. Right? We're doing it actually

three different shuffles. Um so we're

repeating our kfold three times uh for

every now is the thing is we're doing

that for every evaluation. So it is

going to be more expensive than just a

basic kfold.

So we have three different K-fold trials

that we're doing essentially.

Okay, hopefully that makes sense. This

is the repeated K-fold. We haven't

really seen that before. We've only

worked with the Kfold, which would get

rid of this repeats option and only have

uh 10 splits in a random state for the

for the single shuffle that we do. So we

can um recreate that same shuffle every

time. Um but now we're actually going to

do three random shuffles, uh three

different trials. So one shuffle creates

the and then create the 10 splits,

evaluate, then go back and do another

shuffle, another new 10 splits. So, one

thing that should be um clear is that we

get different splits every time because

we're going to shuffle once, right?

We're going to shuffle once and generate

our splits

and then we're going to shuffle again,

generate these splits which are going to

be different and then shuffle one more

time for for three different times,

right? And then get get these splits and

then we're going to get 10 metrics here,

10 metrics here, 10 metrics here, and

then average all of those together.

So, it's a bit more just going up extra

above and beyond for a kfold. Okay.

All right. So, here comes the fun of

when you do grid search. Now, the grid

is actually just a dictionary. It's a

Python dictionary where you declare what

your parameters are going to be inside

the dictionary and you set up a range of

values that you're go or a list. It can

be a list. It can be a range

but some declaration of what you are

going to test and evaluate inside of

your grid search. So the grid is

initialized as an empty dictionary.

And then what we do is we say okay in my

grid I want to check different alphas.

So we're going to add a collection of

alphas in here that we're going to test.

So let me make a comment there. We add a

add a range of alphas to test. And this

range is a this is just like the Python

range. Um

this is just like a Python range um uh

operator here where this is going to be

uh every so it's going to be um every

uh value between

zero and one um uh steps with a step

size

of 0.1. So it's going to try a bunch of

different alphas um between uh zero and

0.1

sorry 0 and one stepping by 0.1. So it's

going to try zero.1

2.3 point 4.5 6 right all the way up to

one.

So that's what this will do and it's a

numpy range. So it's just all those

decimals between 0 to one.

It you can use either that's valid.

Yeah, you can do you can do that to

create a dictionary or you can use the

keyword um dict. You can use either one.

Either one works.

Whatever whatever you want to use.

They're the same.

Yeah. The the reason people prefer

dictionary is because um sets are

created with the same braces.

So it it makes it clear what you're

creating as a dictionary. If you use if

you use this, that's the only advantage

is it's just plainly obvious what you're

making. Uh because technically you can

make a set with the curly braces as

well.

Yeah,

no worries. Um okay, so we have our

alphas here. So what I want you to

notice is that we are going to try out

different alphas and we are that's the

only parameter we are going to try in

our ridge regression. So we're going to

we're going to try ridge but just try

different alphas in the in this range um

in our grid search. So the grid search

CV takes in a model. It takes in our

grid dictionary which is really

critical. We need that dictionary to

declare what we're going to try.

um we need a scoring to say to find the

best. Now remember it uses the negative

to find the lowest which is going to be

the the least negative option.

Um otherwise it wouldn't um just based

on the optimization it would look for

the highest value. Um so the highest

would be closest to zero in this

situation. Um so we use negative and

again we could use squared error. It's

using absolute. We could use um squared

uh either either one works.

Um more typical would probably be

squared error, but um absolute is fine.

Here's where we have our repeated kfold.

So we pass in our um how we're doing CV.

That can be it can be a kfold object. It

can actually just be an integer, which

is say I just want to do 10 splits or

five splits um to to do every

evaluation. But these are the bare

minimum that you need just really the

model and the grid and your CV. Um what

metric you're using to evaluate what's

going to be the best. And then this end

jobs is to parallelize. If you have it

set to minus one, it's going to it's

going to try out all the grid options in

parallel. Um which is nice. It's going

to help speed up the overall search.

Okay. So let me mark that down as n

jobs equals minus one.

tries out the combos in parallel.

So in this situation, we actually don't

have more than one parameter. We only

have the alpha. So we're really just

going to be systematically working our

way through every alpha and evaluating

which one's the best right with this.

And notice that in order to use this

grid search, all we have to do is call

search.fit. So it works kind of like

every other model does, right? It's the

grid search.fit.

And we pass in our data.

And we um once we're once this prints

out the results, you get a results

object

um which has a best score and then a

dictionary with your best parameters.

So, whatever your best grid member was

or grid members, um it prints that out

and you can So, for from that, we can um

grab our best alpha, which it which

let's confirm what that ends up being.

Oops. We need to import repeated kfold.

So we'll import that.

Oh, I didn't. Let's do from

sklearn.linear

model import ridge.

Okay.

Okay. So, it completed the search and

what we found is this is the best score

is 238 for the mean absolute error and

the best alpha that we got was 0.9. So,

the best alpha that worked here, the one

that gave us the best score was actually

0.9 as the alpha. So what it did is it

tried out everything between this range

and 0.9 was the best. So it did cross

validation tried out every single combo

in our grid.

So if we want we could actually print

out

print our grid so we can see

what our combinations were.

So, it tried out all of these guys and

the best one that we had was 0.9.

Okay, so pretty cool how that works. And

if we had other parameters, like if we

were doing a elastic net, we could add

those into our dictionary and it would

do all combinations of those. So if we

did um so for instance to add to our

grid we could do grid

um L1 ratio

this would be for like an elastic net

right now the ridge regression by itself

doesn't have an L1 ratio parameter but

just as an example um we could try out

different ranges

um similar range different one um maybe

an exact list whatever we want to do. So

this is going to try out different ones

between 0 to one as well. And so it's

going to try out every combination of

these from this grid.

Okay, if we did that. But again, this

the ridge regression doesn't have an L1

ratio. The elastic net does. So that the

ridge regression only has an alpha to to

as a hyperparameter. So we're only

testing out that one.

Okay, so that's grid search CV.

Pretty useful. This is pretty useful in

doing parameter tuning again when you

want to try out ranges of different

values and you can evaluate those to see

which one is your best and then we can

use that best going forward. So we can

for instance this is what this code does

below it is it fetches the best. Um you

can do it this way or you can do it um

the alternative is to do results.b best

params

and then you can just grab it like this

alpha.

Either way you can do get or like this

um and it this is just a dictionary,

right? And you can grab your alpha. So

that's the 0.9 um and we can pass that

alpha into the ridge regression and go

back and refit it to our data um and

then use that model going forward. So

the grid search really just evaluates

those different options, tells you

what's the best according to this score,

right?

And you should, by the way, you should

interpret this score in the positive

sense. It's only negative because we're

purposely making it negative to find out

what the lowest option is, right?

Because the lower is the better. So we

we purposely make it negative to make it

whatever is the least negative is the

winner. Um more negative is worse.

So it's the really positive version of

it is the is the true result for the

error. Um and they are a tool from

scikitlearn to put together your model

with your pre-processing steps. So they

kind of get automated together. Um and

they combine everything into kind of a

streamline process. You're going to see

what that looks like, but it's a really

nice um feature of scikitlearn. Um why

would we care about pipelines? They help

organize our code um so that we ensure

that we basically always run the

pre-processing steps before we train and

use a model to with the predictions. Um,

so it bundles those steps together,

minimizes the risk of forgetting a step

because one of the things that can

happen is when you do pre-processing, if

you're doing it on the training set, you

have to do it on new test data as well

when you put it through your model

because your model is training against

that pre-processed data.

So in order to make sure you never

forget that, you can bundle it all

together in a pipeline which is going to

make things really really easy to use

and and make sure that those steps

happen in a repeatable way. Um and it

makes things easier to uh deploy that

model as well because everything is

together in one pipeline. So in the in

the industry, I've seen this a lot. Um

you know, people will do their initial

exploration steps and initial model

building. They may not use pipelines

right away, but as they found their

model, um they'll generally move it into

a pipeline and all their steps into a

pipeline so that it's uh easier to work

with um when you're when you're

deploying it and actually using it uh in

in the real world. Um, so here's what a

pipeline generally looks like. It's from

scikitlearn. It's this pipeline object.

Um, and the pipeline is made up of steps

that we're going to see that that are

various um uh basically um kinds of

pre-processing we've seen before like a

scaler or um filling in missing values.

Those kind of things we can put here in

the steps which is basically a list. um

steps is just going to be a list of

scikitlearn functions that we can apply

to data. One of those being a model. Um

and then whenever we use the pipeline,

it's basically um you know it's going to

be something like pipeline.fit

or pipeline.predict.

So the pipeline kind of behaves like a

model. It's just going to contain many

more steps than that like the

pre-processing steps we've worked with

before. Um, and it also has some

capabilities for caching. So you can

like uh cache some of the data in

memory. Um, so that if you're reusing

the predictions, it kind of goes faster.

Um, so there's some options for that

too. I'm not too concerned about that at

this stage, but the main thing is going

to be filling out our steps and then

using the pipeline.

Okay.

Um, so some important bits of

information about the pipeline is that

it is going to be a sequence of data

transformations that will have at the

very end of the pipeline the model

because of course we're going to do

transformations and then train a model

or predict with a model. So every

Oh, can you guys hear me? Okay,

not able to hear me. Thanks for letting

me know. Can you guys were you able to

hear me so far?

Okay. Make sure. Yeah, it might be on

your internet or your your uh Yeah, it

seems like seems like it's good. So,

no, you can't hear me. Check your

volume. Check your headphones if you're

wearing headphones. Oh, no issues. Okay,

perfect.

Okay. Yeah, local internet issue. Yeah.

Okay.

Always let me know. always let me know

because it could be the case that it is

me. So, um always always make sure to

let me know. Um but sounds like yeah,

you may want to check on that. Um

so, okay, what I was saying is every

pipeline is going to have a uh a

sequence of steps that go first and then

the model at the end. Um so, the order

really matters. Um

uh so the order matters in the sense

that we want our transformations to go

first. Things like scaling, things like

filling in missing values, we want those

to be first and then we want our uh

model to be last because we want those

transformations to happen prior to

training or prior to prediction. So

usually what you'll see in these

pipelines is a model at the end, right?

a model that's going to be at the end of

the pipeline because we want basically

our processing steps then our training

or processing steps then our

predictions. Um so everything in the

pipeline though is going to be from

scikitlearn. Uh that's how it gets

automated in the sense that all of those

things are going to have fit and

transform functions built into them so

the pipeline can use them. Uh, and then

the last step is going to be a model

that has a fit and a predict. So it's

pretty standard that the last part of

the pipeline is just going to be a

model. Um,

uh, so we can um, as we do more

modeling, we're going to play around

with the pipelines quite a bit and see

how we can change up some of the

parameters. like if we want to change a

model's parameter um we can actually

adjust it to do things like uh grid

search or cross validation. So um we're

going to see some examples of some

pipelines but for right now mostly what

we're going to see is how to build one

and then how to use one. And then as we

get into lesson four, we'll get some

more practice with pipelines because

we're going to start using them quite a

bit uh to build our models rather than

do manual steps uh all the manual

pre-processing

um and then kind of building a model

from there. We'll just include all of it

together in a pipeline.

Okay, so the example we're going to do

is with this housing with ocean

proximity. So we've actually looked at

this data set before. Um so we have uh

this ocean proximity data set that has

the feature of like how close it is to

the ocean like the bay or the less than

1 hour. Remember we had that and it had

the median house value for different

neighborhoods. Um so we're going to work

with that one again. Let me make sure I

have that one uploaded.

You guys should have this one. It should

be in your uh data sets.

Um, I'll I can upload it here in case

you don't have it though.

Does this use multi-threading? I think

it does. Yeah, I think in order to do it

can do uh um I think it can do

processing in parallel for some of the

pipeline steps. Um, now does it use that

all the time? Not necessarily because

some of it is sequential in nature where

you have to do one step and then you do

the next step and then you do the next

step. So it's not like you can do them

in parallel.

Um in terms of the like you need to know

the output of one step to compute the

the output of the next step. Um so it

can but it it doesn't always lend itself

well. The thing that will use

multi-threading is is like the training

process can be parallelized

like the fit um can be and for some

models it can be parallelized not every

model

it. So long answer is or the short

answer is that it depends

depends on what kind of transforms

you're doing and what kind of model

you're using. if you can really take

advantage of that.

Okay, so we load our data here and take

a look at that. Um, do you guys have

this data set? Are you able to load it

in? If you're following along, are you

able to load it?

Okay.

And and again, we've worked with this

data before, so hopefully it's somewhat

familiar. Remember, every row represents

a neighborhood, and it has a we're going

to end up trying to predict this median

house value as our target um variable,

our dependent variable. Um and we're

going to use the rest of these features.

Remember that um this feature is in

particular going to need to be one hot

encoded,

right? It's going to be one hot encoded

because it is currently a string. and we

need to turn that into a numerical

feature which is the one hot encoded

feature. So we're gonna have to do that

but we're going to do that as part of

our pipeline.

Okay. So we'll be able to include that

in our pipeline steps uh to to do one

hot encoding which is nice.

All right. So we're going to split apart

our data um as we normally do. So we're

going to uh create our feature uh data

frame which is everything but this

median house value. So we go ahead and

drop that column and then our target is

the median house value. So it is just

that column here. Pretty standard. Um

and then we're going to train test split

and um split it into 30%

uh test data. And again random state you

can choose whatever you want to be. that

just affects the shuffling. Um, so

whatever doesn't really matter what it

is. It's just so that when you rerun

this, you get the same result in the in

the shuffle.

Okay, so we have our train and our test.

So you want to make sure you run those.

All right, so what we're going to do is

take a look at our data

and see if we have any null values. Um,

if you guys remember this data actually

did have null values. You can see it

here in this this guy and exactly how

many there are is from this the sum. So

we have um 162 nles in in this data. Uh

and this is just a training data. So of

course you know the test data could have

that in there as well. Um so that's

something we're going to want to make

sure we fill in the blanks on any data

set we use whether we're using the

training or test set. Um, like if we're

doing training, we want to make sure

that gets filled in. If we're doing

predictions with the test set, want to

make sure that gets filled in. Um, so we

we should be doing that. Um, now

what we're going to do is use this data

to help uh train our pipeline or or use

with our pipeline. We need to construct

our pipeline. So far, we've just split

apart our data. We haven't done anything

with our processing steps in our model

yet. Um so roughly

it this should be the flow of our

pipeline. What should happen is we

should be doing some type of feature

scaling

um some type of uh feature um

manipulation. So that could be

engineering, that could be um that could

be uh doing the one hot encoding. Um so

extracting new features like one hot

encoding,

one hot encoding. Um we are going to be

doing that and and by the way this is

split up into this is when we use our

pipeline for training.

Um it's going to look like this where we

do our scaling, we do one hot encoding.

Um we have our model here. Um so that

could be a linear regression, that could

be a lasso, that could be a ridge, it

could be elastic net. Whatever model we

end up using is going to be last in the

pipeline. And we're going to run this

pipeline. Ultimately, we're going to run

pipeline.fit,

right? We're going to run a fit

function. and we get a fitted model as

the result of this pipeline.

Then when we use it when we use our

model for prediction,

we use our model for prediction in this

lower part, it's the same pipeline, same

exact pipeline, but it's this model has

now been trained.

So we now have a trained model here. So

the great thing about the pipeline is

it's the same this is the same pipeline

that we're using here. So it's just

going to it's going to repeat those same

transformations. It's going to do our

scaling. It's going to do our one hot

encoding. It's going to use our model

and it's going to generate predictions

and generate uh we can we can do

predictions. We can do evaluation like

an across validation. Um we can use it

however we want to use it. Uh but notice

that the pipeline makes it consistent

between training and test. We're using

the exact same transformations

and the model is last. It's it's either

being trained or it's being used for

prediction but it's last. Our

transformations are up front which are

things like our scaling, things like our

one hot encoding, right? Those happen

first. No matter what data we put

through there, we put our training data

through there, we put our test data

through there, they're going to go

through the same steps,

right?

So that's that's the design of the

pipeline. That's what it's supposed to

do. So our job is to create those steps.

So we need to create those relevant

steps and then put them together into

this pipeline. Okay, so that's going to

be the code we're going to see coming up

is we're going to build out these steps

and then put them together into the

pipeline.

Um, any questions on this diagram? Does

it make sense what we're trying to do

with this pipeline? We want to have

repeatable steps during the training,

during a prediction process.

Okay.

All right.

All right. So, um, a couple of things

we're going to need is, uh, to first of

all, let's jot down what steps we're

actually going to do. We're going to

need to deal with missing values. So,

we're going to fill in we're going to

need a pre-process pre-processing step

that fills in any nles. We always need

that, right? So, if there's nles, we're

going to fill them in somehow.

We're going to define how we do that in

our in our step. Um, and we also need to

one hot encode. And we need to scale,

right? Those are pretty standard steps

that we've dealt with whenever we're

building these models, right? So, pretty

standard things. fill in nles one hot

encode any categorical data whatever

however much we have and then go ahead

and um standardize which is the scaling.

So this this just is the same word for

scaling our numeric features. So we're

going to we're going to define those. Um

so that's why we're going to go ahead

and import from pre-processing. We're

going to import our scaler. Um again we

could use minmax scaler here. We're

going to use standard scaler. Um but we

could use minmax. Um we have our oneh

hot encoder here. Now usually when we do

oneh hot encoding we use pd.get dummies.

This does the same thing as that. But

because we're going to be building a

pipeline we actually want the

scikitlearn version of git dummies. So

this is the scikitlearn version of git

dummies here. And it and we have to use

that version in the pipeline because

everything in the pipeline needs to be

an sklearn object. It needs to be an

sklearn tool or object.

So um instead of using pandas get

dummies, we're using one hot encoder

which is does the same thing. Okay. In

fact, it just this basically just uses

pd.get dummies um under the hood.

Okay. So, it just uses that. Uh,

anyways, it's just code that builds on

builds on that.

Now, what's really nice here is we're

also going to use from sklearn.impute,

we're going to use a simple imper. Now,

what this is is an automated way to fill

in missing values. So, this is a fancy

way of basically doing the the fill na

on a data frame. So, simple imputer um

we are going to basically fill in the

blanks. What we're going to do when we

create this object is give it a strategy

of how to fill in blanks. Should you use

the average? Should you use the median?

Should you use the max? Should use the

men? Should you use a default value?

We're going to tell it what to do in

this object.

Okay. So, we're going to we're so we're

going to use this as our automated tool

for filling in missing values. So,

that's really nice. it has. So this is

going to be a critical part of our

pipeline an imputer that's going to fill

in missing values.

So we have that

yeah coding to reduce coding. Exactly.

Uh we have our pipeline now. So we have

our pipeline. So our pipeline is going

to hold everything. So we need the

pipeline object um to hold everything

and that comes from sklearn.pipeline.

Um, so everything's going to actually go

into a pipeline object. We're going to

see how that looks. Um, and finally,

we're going to from skarn.compose, we're

going to use a column transformer. The

reason we're going to do this is because

we are going to specify for some columns

like the numerical features, we should

be scaling.

For some columns like the categorical

features, we should be one hot encoding.

So the column transformer will allow us

to map different transformations to

different sections of columns which is

really useful. So this is actually going

to be a critical part of our pipeline to

apply to make sure we only apply this to

numerical features and only apply this

to categorical features. Right? So this

column transformer will help us um to to

apply pre-processing to particular

columns. Um like that ocean proximity is

the only one that really needs this but

every other column is going to need this

all the numerical features.

So we're going to use this column

transformer and again we're going to see

how this looks but just trying to give

you an idea of why we're importing all

these things.

Okay, so let's import those.

Uh, this mentions about the column

transformer. We just talked about it. It

allows us to have a particular column or

group of columns get the right

transformation. So again, uh, looking

ahead to our pipeline, the numerical

features are the ones that are going to

need scaling, but the categorical

features are the ones that are going to

need one hot encoding. However many

categoricals there are, in this case,

there's really only one, which is that

ocean proximity. to go back to our data.

Um, you can even see that in the info,

there's just that one. Um, and we see

that here, right? Just this one string

column that should be one hot encoded.

All these other guys should be scaled,

right? They should all be uh uh standard

scaled.

So, this will allow us to specify those

distinctions.

All right. So, let's get started

building our pipeline. So, this is going

to be really cool. We're going to build

out the pipeline. Um, let's extract our

numerical data and our categorical data.

Now, this is a really neat way of doing

that that I'm not sure we've seen

before. Um, so what this does is we'll

take our data frame, particularly our

training data frame, and select our

data.

That's what this select dtypes does is

select data from it. Um, which includes

only the object type columns. So only

the object types. Now what's that? The

object type is the string, right? So

this should select only this column

because it's in the include.

We go here include only object types in

the result. And so this should only have

our one categorical column which is the

ocean proximity. So housing cat is going

to have a reference to our uh it's going

to be a list that has a a basically just

our ocean proximity feature because this

select dtypes will make sure we only

pick object types and um

grab those columns. So this is a way to

neatly grab um our categorical features

here by including the object types. Now

on the flip side we can exclude object

types and get everything else. So this

is going to be all other columns which

is excluding the object. So this is

excluding this meaning we should get all

of our numerical features that way. So

this will be all of our numericals

by excluding the object type and this

will be our housing num which is short

for numerical. So this excludes

the uh object type

meaning all numerical

features

right all numerical features there.

Okay.

So, if we were to uh let's double check

this. Let's sanity check this. If we

were to print out the housing

cat, um this should be just the ocean

proximity feature, which it is. So, just

that one. If we were to print out the

housing num, this should be all the

numerical features, which are all these

guys. So it's just a reference to those

columns so that we can uh use those

later when we're mapping uh this

transform needs to go to this column

like the one hot encoding needs to go to

this column and the scaling needs to go

to these columns right so we have those

uh names of those columns already at our

disposal. So, we're just doing that.

And this is just a

simple check.

Uh, are you guys able to run this?

If you're following along, let me pause

there. Make sure I'm not going too fast.

Uh it so the the issue with a specific

data type like that is none of these are

ants. They're actually all floats. So we

did float. I think that should work. But

yes, that's the idea.

Great. I'm glad to hear that right there

with me. Great. Glad to hear that.

Okay. So, we have our columns picked out

here, which we're going to use later.

Okay.

All right. So, let's go ahead and build

out our steps for each of these types.

So, um for our numerical features, let's

build out our pipeline steps. So what

we're going to do is build out a

numerical pipeline. And it's going to be

a pipeline with a list

of tupils. And the reason these are

tupils is because every tupil has a

name. So here this is a name that we can

it can be whatever we want it to be. So

we're calling it imputer. We could call

it anything we want. We could call it

fill in the blanks. We could call it

null filling. Call it whatever you want.

We're calling it imputer because it's

that's a pretty um easy name for it. An

accurate name to what it's doing. Um but

the important thing is after the name

you give it, you put in the scikitlearn

object that you are going to use to

operate on your data. So in this case,

we're using a simple impery

of median. Now that's a choice. We could

use a strategy of mean, max.

Um, we could provide it a constant

default value. But what this means is we

are going to fill any blanks we find in

those columns with the median value of

that column. That's the strategy for the

imper. So that's pretty cool. This is

kind of an automated way to fill in the

blanks using for any column using its

median,

right? And so we could change that. We

could put mean here or max or min or

whatever. Um

but we are filling in the blank on any

column with its median. And the reason

this works is because we are going to

apply this pipeline only to these

numerical features. So that is fine.

We're we're not going to apply it to the

categorical features. We're going to

apply it to only those numerical. So it

should have a median value, right? So

that that's totally fine. So we're going

to now look at how we're constructing

the steps. We have a list of tupils.

Here's one tupil

which is the imper with a simple imper

of strategy median. And then we can have

as many tupils as we want which

represent processing steps. So every let

me write that down. Every tupil

represents

a pre-processing

step on our data.

Okay, so we have an imputer step named

imputer and the reason it has a name is

just so you can reference it in the

pipeline if you need to. So you so it

has like a a reference name um that you

give it. Um but this is the more

important part is the actual scikitlearn

object that's doing the processing. So

in this case a simple computer but

notice that we have a secondary step

which is our scaling. Now this makes

sense. This is something we should be

doing to our features is we should be

scaling them. So here we we say okay

let's fill in any blanks first.

By the way order

matters.

So, and what I mean by that is the

simple imper

is before the scaler. Now, that's

important because what that means is we

should be filling in any blanks before

we attempt scaling.

So, that order actually matters. We're

going to fill in blanks first in this

list. That's first. We're going to fill

in blanks. Then we are going to scale

right then we scale which makes sense

right so we we fill in blanks first then

we apply the scaler to scale our

features so those are our two steps

so so pretty simple um we are building

out our two steps now this is just one

piece of the puzzle we are going to put

this pipeline together with our one hot

encoding that's going to be coming up

next and build out our final pipeline.

But this is um a a pipeline that has two

steps that will actually be used with a

larger pipeline coming up where we we do

one hot encoding to our categoricals and

then we put a model in there at the end

to train and and use for prediction. So

um pipelines can actually be composed is

is uh something to realize there is that

we can have a pipeline that contains a

few steps. We can have another pipeline

over here that contains a few steps and

we can actually um kind of put them

together into a final pipeline that has

both pipelines uh kind of merged

together. Okay. So we're going to see

that coming up when we construct our

final one. Our final one, as you can

imagine, needs to handle this mapping of

basically saying, let's do one hot

encoding to these guys and then do this

pipeline here to these numerical

features. That's what our final pipeline

needs to handle, and it will. We're

going to build that out.

But let me pause here. Um, were you guys

able to run this? Are you with me on

this this pipeline here?

Does that make sense? Those two steps

one is filling in blanks with a median

whatever column. So where so this is

this is what's so amazing about this is

this is going to automatically search

for nles and if you come across a column

with a null, it's going to use the

median of that column

to fill in the blank, right? To fill in

those nles.

Okay,

great. Glad to hear. Glad to hear.

Okay.

All right. So we are going to now um put

this together with a column transformer

to basically say what steps are going to

be mapped to what columns.

Um so now you can see what we're doing

here is using the column transformer

which is going to be a list of tupils

again. So this is another um list of

tupils.

But the important thing is um

each tupil

has a name

followed by so it has a name uh which

again is is generic. You can say

whatever you want it to be. So here

we're kind of shortening this to

numerical. This is short for

categorical. But the important thing is

it's followed by a pipeline

slashstep

followed by a pipeline slashstep

um followed by a uh followed by a list

of columns that it applies to. So you

can see that pattern here. What we're

saying is we're going to apply that

numerical pipeline we just defined. So

this is saved in a numerical pipeline

object here. We're going to apply that

to those numerical features. So this is

that list

of numerical features here. So that's

how we do the mapping. We have a tupil

here that says okay apply these steps to

these columns.

Those go together in that tupole, right?

Apply these steps to this uh these

columns. And then apply this step. Now

what is the step? This is a one hot

encoder

which is going to uh uh encode um those

features and it's going to uh ignore um

basically nles for now. That's a choice

but it's going to ignore um uh basically

ignore nles and and uh skip over them

for now. We now we know there's no NLES

because we already did an is NA from

before and we know there's not any NLES

in that ocean proximity. So this isn't

going to be an issue. But that's what

that would do.

But we have a one hot encoder here which

we're going to apply to our categorical

features. Now of course that's just the

ocean proximity feature but that but

again you see the pattern in the tupil

is apply this transform which is a one

hot encoding to this column apply these

numerical transforms which is a whole

pipeline. So it's two steps in a

pipeline of um

uh an imputer and a scaler are going to

be applied to this

really nice. So those are going to be

all together in this column transformer.

And that is our way to signal that for

these numerical features, use these

steps. For our categorical features, use

this step. And and you know, if we had

more than one step, we were applying to

categorical. We could build a pipeline

for the categorical and it would and do

the same thing. We have more than one

step here. And so it's good practice

when you have more than one step to just

put that in a pipeline because we have

more than one step. We'll just put that

in this list inside of the pipeline and

we can map that pipeline to those

features. Here we only have one step. So

it's okay to just put that there um and

apply that to the categorical features.

But if we had more than one step um it

would be good practice to put that in a

pipeline

which is what we do here. Right? This

pipeline is being mapped to these

features. This step is being applied to

this feature.

Okay,

how about that? Are you guys able to run

that one? Does that make sense what we

have set up so far? So, we're almost

there. We almost have our final

pipeline. We have our pre-processing

basically done to say our numerical

features should be processed with that

other pipeline and our categorical

features should be one hot encoded.

We're getting close. The only thing

we're really missing here is a model.

The only thing we're really missing is

to have our final model training

pipeline is to actually include a model

which should come at the end.

Right? So it should we should be doing

these steps first

then doing modeling which we know right

we we've done that uh many times. We've

done our pre-processing and then we do

our modeling.

Any questions on that?

Okay.

Fantastic.

All right.

So, if we wanted to uh see if we wanted

to test this so far, um we could. So, we

could run the pre-processing and

actually run a fit transform on our data

and this will um basically apply that

pipeline to the data. Now, this would be

a sanity check. This is a good This is a

good kind of um This is a good sanity

check that our pre-processing

works. So, it's doing what we expected

to do. It's not our final pipeline

because we don't have our model in there

yet, but this is just to ensure that all

of the features are kind of behaving as

we expect. So, we can uh we can do that

and we can take a look at the um

results. This looks pretty good. this

all of our numerical features ended up

scaled

which is pretty good and we have one hot

encoded features for that ocean

proximity over here.

Okay, so this looks pretty this looks

reasonable of those steps being applied

to the right columns. But this is a good

kind of sanity check to just run our fit

transform on our data to ensure those

steps are actually happening and they

are. You can see here the result of the

scaling and the uh the one hot encoding.

So that that all looks pretty

reasonable,

right? And uh what we should also do is

make sure there are no nulls in this

which there shouldn't be because we did

the imputer. So we should be doing uh is

na dot

sum

And there is no NLES anymore. So that

looks pretty good, right? Those got

filled in uh by doing our steps. Our

pipeline steps executed really nicely on

our training data. Um and and we were

off and running. And there's nothing

unique about the training data. We could

do this to our test data as well

and verify that those steps are running

and they would, right? There's nothing

really that special about running it on

the training data. Um, it should also

work on the test features as well, and

it does. You can check that for

yourself.

Okay.

All right. So, that's pretty cool. We

can uh verify all that's working.

Any questions on that?

We're almost there with our full

pipeline. This this is this is not the

full pipeline, but this is something

that will run during our full pipeline.

Of course, our features are going to be

transformed according to those steps and

then it will be uh put into our model to

either predict or train with. Um

so let's do that. Let's actually build

out our final uh model here. So it's

actually going to be really easy to do.

All we need to do is um put in our

model. So here we're going to import the

ridge model here. Now we could use any

we could use linear regression, we could

use lasso, we could use elastic net. Um

we're just going to use ridge um uh um

just to test it out. And um we are going

to uh now put in a final pipeline. So

we're going to use our pipeline. And so

we're going to create a new one here.

and map our pre-processing to our

pre-processing that we've already built.

So, this is a column transformer that

already has all of our steps. And then

notice what comes after it is just the

model. Now, that's pretty pretty basic,

but it makes sense that it should come

after that model. Um, and of course,

this is a generic name. We could we can

name it whatever we want to. um model

ridge is pretty reasonable um to because

it is a ridge uh regression but uh of

course we could we could change that.

Okay, so that builds out our uh final um

pipeline. So now we have a pipeline and

what's great about that is this signals

that all of these steps should be

completed prior to doing anything with

this model. So all of those processing

steps are going to run and then we're

going to do ffit orpredict and that so

that's really great. It ensures that all

those steps are running together every

single time we call predict with this

with this model. So we're just going to

use the pipeline in place of the model

to ensure that all of those steps are

running together. And this is our this

is kind of our final pipeline that we

would use uh with like something like

ffit or predict.

So let me make that uh a note of that.

Now we can use this final pipeline just

like a regular model i.e. pipeline.fit

or pipeline.predict.

We could use it in ei in either fashion

uh to to train the pipeline would be

this guy and then use the pipeline to

predict would be this. And what we

should realize is under the hood these

steps are running first and then we

train it or these steps run first then

we use it for prediction.

Okay.

Questions on that? Does that make sense?

On this final pipeline here, it's just

now it it's really cool because we have

a pipeline

made up of a of a pipeline really,

right? A pipeline made up of a pipeline.

But that's scikitlearn allows you to do

that to compose pipelines in this way.

That's that's pretty uh pretty uh normal

there.

Okay,

what I want to show you is we can

actually use this pipeline in a grid

search. So that's pretty amazing. We can

use this pipeline in any way we can use

a mo like a regular model. It's just

that now our pre-processing steps have

kind of been packaged together with our

model to ensure that they always run

anytime we do any processing with this

model. Um so for instance we can do a

grid search just like we did with a

regular with with just a model right

with just this. Um we can do the same

thing with the whole pipeline. Um, so

the only catch is that you want to make

sure in your grid you name things in the

appropriate way inside of your your uh

keys in your dictionary. So uh for

instance um inside of the grid uh we're

going to set up the alpha that would be

used with this ridge regression by

referencing its name. So this is model

ridge is this is the name of the model

inside of the pipeline. So you want to

make sure that goes first.

And then what scikitlearn does is it

recognizes parameters that belong with

this model by using a double underscore.

So the so you have underscore underscore

alpha um here. So the double

uh underscore

signals a parameter

belonging to model ridge. in the in the

pipeline.

Okay, so we have a model ridge is just a

reference to the model in our pipeline.

That's the one we're going to test out

these parameters with. And

underscore_pha is just a way to say this

alpha belongs to this model. Okay, it

belongs so it's going to be used with

that model in our pipeline. Um otherwise

it's going to work exactly the same way.

It's just we need to line up this naming

convention of of scikitlearn.

You just have to reference this to

whatever name you provided here and then

underscore parameter. So L1 ratio alpha

whatever right would go there.

Okay. So there is a range from 0.1 to

two uh step size of 0.1

um and then we do our grid search CV. So

this is exactly the same setup as we had

before. It's just that our model is now

the pipeline. So our pipeline is going

in there. Um we have our grid going in

there. We have our scoring is the same,

you know, negative absolute error. Um

we're using five-fold cross validation

and we're parallelizing that search. Um,

so we're going to search through these

alphas and uh basically fit this to our

um data and find the best um find the

best alpha.

So it's going to try out all those

combinations and try to come up with the

best alpha.

So looks like the best alpha was 0.1 for

the ridge.

Okay, is the best. So then um if we

wanted to we could uh then predict using

the model um which would be doing

something like this. Um and we could

also go back and do something like so we

could

now use um this param. So we could do

model

um equals ridge

and then we could put in our alpha

um alpha is our results our best

parameters and then we get that model

ridge alpha and then we just rebuild our

our pipeline

equals um pipeline and then we uh put in

this new model here. So we could do

this. This would be going back and just

um putting in our best alpha here for

this model and then uh ensuring that's

part of our our pipeline. So we're just

overwriting that pipeline with the best

model there

to get the best model in our pipeline.

Okay,

so that's all this is doing is just

initializing a new um let me actually I

can put this code in here.

This is actually just getting this is

just getting a model with the best alpha

and then reinserting that into our our

uh we're just overwriting our final

pipeline there with the best model that

we have.

So pretty cool that pipeline can be used

basically exactly like a model, right?

It's it's going right here in the grid

search and being used uh entirely like a

basic model. So we do ffit

um and that allows us to use it. We

could dopredict. We could even do

pipeline.predict once we we could go

back and do final pipeline.fit

um with this and then final

pipeline.predict with this and evaluate

Okay,

so pretty cool that pipeline can be used

uh basically exactly like how a model

would be any way we' use a model.fit

model.predict, we can use a pipeline.

So grid search is for instance something

that can use a model in there. Um but

instead of just a model, we're ensuring

we have our pre-processing steps kind of

bundled with that model in this

pipeline.

Any

questions on

uh this example so far?

Were you guys able to run it up to here?

Were you able to run the grid search?

Okay, great.

Okay.

Okay. So, this is this is uh just

showing you what's actually happening

underneath the hood is uh you know,

we're doing some scaling. We're doing

some one hot encoding um

and we're doing some uh we're doing a

model here. And that's all part of our

pipeline. Um, and then we can use the

pipeline however we want. So for

example, I know it's not here, but for

an example, we could use um once we do

once we have this final pipeline um we

can can use the final um

pipeline to predict. So we can do um

predictions

equals final

pipeline.predict

and then we can pass in our test data.

Now what happens on this is once we have

ran our our pipeline.fit we have a

trained pipeline and then when we run

this final pipeline.predict uh this data

is going to be transformed.

It's going to go through those

transformation steps and then we would

apply our model to it at the end uh to

to make those predictions and then we

can evaluate those predictions which is

what we're doing kind of here.

Right?

Okay.

All right. So in conclusion uh we have

gone through a lot of stuff here. Um,

we've gone through regression, we've

done the regularization on regression.

So hopefully we have a good foundation

on regression. Um, what we're going to

do in a little bit is actually do some

additional practice with regression on a

new problem. We're going to do a

capstone problem and do some additional

regression work with that. Um, so we'll

do that next. Um but the other thing we

learned is how to evaluate the

regression using things like mean

squared error, RMSSE which is square

root of that. Um which is which is

really cool. So we have a sense of that

error which is our distance from our

prediction to the actual value. That's

always what these uh that's always what

these things are doing like this, right?

This mean absolute error metric from

scikitlearn is computing the average

distance from these predictions to these

test labels that we have, right? Those

actual values. Um, and that gives us a

sense of on average how far away are our

predictions

um to see how good of a model that we

have, right? And we should be evaluating

that error generally

um against the scale of our targets to

see, you know,

uh how far off we typically are.

Okay. Any questions at all on this

lesson on regression? Uh anything we

covered up to this point? We're going to

do some more practice with the next

we'll do the capstone. So we get so we

just do some more regression problems.

Yeah, it's a that's another bad score.

It's a little bit hard to interpret this

though because it's m ae. Um, so one

thing we could do is is compute mean

squared error and then take the square

root of it to get the RMSSE which is a

much better uh evaluation metric in

terms of our target. Um so we could

actually run that. Uh if we go back here

and um we could generate for instance we

could generate the MSE which is the mean

squared

error

and it's it's the same exact function uh

of using our predictions.

Um

and then we could just print that out.

Mean squared error.

So we have mean squared error and then

what we can do is let's take the um MP.

square root of that.

So that way we can generate the RMSSE.

So yeah that I mean that's pretty bad.

That's uh pretty bad. Uh now let's let's

go back and look at our

uh data though. So let's take a look at

the average for our y. Um remember one

thing we should be doing is taking a

look at um what our uh let's take a look

at y test mean

to get an average value. So the average

value is in the 200,000s. So

this isn't this isn't awful. This is

70,000. It's still a decent amount of

error. It's not as bad as the models we

have before though, right? This is an

average median price of the house is in

the 206,000 range and our error is off

by like 70,000,

right?

So, it's not good. Um, but it's not

hor like as bad as the it's not as

horrible as we've seen so far. Right.

This is a little bit better of a model.

A little bit better. closer to zero

would be better, right? Um but the

smaller the better. But uh remember this

is the um these even the mean absolute

error is is technically in similar units

as the as the uh um

as the target. So 50,000 60,000 here

70,000 it's still a decent amount of

error in terms of 200,000.

Uh so far we only come up with models

and test their accuracy with available

data. We haven't used a model to make

completely new predictions on No, we

haven't done that. Uh except we know how

to do that. Um it would so to make

predictions on new data would be exactly

how we're making them on our available

data because we actually do that all the

time. If we go back down to our model

building,

um it's it looks just like this, right?

where we take so for instance we do

predictions all the time on test data

that was never involved in the training.

So it's it's as if this data mimics new

data that we've never seen before. So if

we had new raw data it would just it

would be the same exact process. the new

now with our pipeline it makes it a

little bit easier because with the

pipeline

um the raw data will go through those

transformations which it should right

the raw data should because if it's

missing data it needs to be filled in if

it has categoricals it needs to be one

hot encoded so that's the purpose of the

pipeline actually is to make sure that

if we're dealing with raw data um those

steps can happen on the data before it

goes into the model. Right?

So we so

that's kind of the purpose of the

pipeline

is to ensure that we run those steps

ahead of using it using a model with it.

But but ultimately that's how it uh any

scikitlearn model is going to be doing

the predict even if it's a pipeline

right it's going to be uh we just go

back down here it's going to be um

predict it's always going to be that on

new data

>> hello everyone in this session we will

cover all the important supervised and

unsupervised learning algorithms that

are widely used with hands-on

demonstrations in Python our instructors

with rich experience in machine learning

will take us through this course but

before Before we begin, make sure to

subscribe to the SimplyLearn channel and

hit the bell icon to never miss an

update.

So, we will start by understanding the

basics of machine learning from a short

animated video followed by the

difference between supervised and

unsupervised learning. We will then jump

into learning the various algorithms

from scratch. So, we will understand

linear regression, logistic regression,

decision tree, and random forest. We'll

then look at support vector machines and

K nearest neighbor algorithm with a

hands-on demonstration in Python.

Finally, we'll get an idea about

unsupervised learning algorithms such as

K means clustering and principal

component analysis. We will conclude

this session with regularization in

machine learning. So let's get started.

>> We know humans learn from their past

experiences and machines follow

instructions given by humans.

But what if humans can train the

machines to learn from their past data

and do what humans can do and much

faster? Well, that's called machine

learning. But it's a lot more than just

learning. It's also about understanding

and reasoning. So today we will learn

about the basics of machine learning. So

that's Paul. He loves listening to new

songs.

He either likes them or dislikes them.

Paul decides this on the basis of the

song's tempo, genre, intensity, and the

gender of voice. For simplicity, let's

just use tempo and intensity for now.

So, here tempo is on the x-axis, ranging

from relaxed to fast, whereas intensity

is on the y-axis, ranging from light to

soaring. We see that Paul likes the song

with fast tempo and soaring intensity

while he dislikes the song with relaxed

tempo and light intensity. So now we

know Paul's choices. Let's say Paul

listens to a new song. Let's name it as

song A. Song A has fast tempo and a

soaring intensity. So it lies somewhere

here. Looking at the data, can you guess

whether Paul will like the song or not?

Correct. So Paul likes this song. By

looking at Paul's past choices, we were

able to classify the unknown song very

easily, right? Let's say now Paul

listens to a new song. Let's label it as

song B. So song B lies somewhere here

with medium tempo and medium intensity.

Neither relaxed nor fast, neither light

nor soaring. Now, can you guess whether

Paul likes it or not? Not able to guess

whether Paul will like it or dislike it.

Are the choices unclear? Correct. We

could easily classify song A. But when

the choice became complicated as in the

case of song B. Yes. And that's where

machine learning comes in. Let's see

how. In the same example for song B, if

we draw a circle around the song B, we

see that there are four votes for like

whereas one vote for dislike. If we go

for the majority votes, we can say that

Paul will definitely like the song.

That's all. This was a basic machine

learning algorithm also. It's called K

nearest neighbors. So this is just a

small example in one of the many machine

learning algorithms quite easy right

believe me it is but what happens when

the choices become complicated as in the

case of song B that's when machine

learning comes in it learns the data

builds the prediction model and when the

new data point comes in it can easily

predict for it more the data better the

model higher will be the accuracy there

are many ways in which the machine

learns it could be either supervised

learning unsupervised learning or

reinforcement learning. Let's first

quickly understand supervised learning.

Suppose your friend gives you 1 million

coins of three different currencies. Say

1 rupee, 1 and 1 dirham. Each coin has

different weights. For example, a coin

of 1 rupee weighs 3 g. 1 euro weighs 7 g

and 1 dirham weighs 4 g. Your model will

predict the currency of the coin. Here

your weight becomes the feature of coins

while currency becomes their label. When

you feed this data to the machine

learning model, it learns which feature

is associated with which label. For

example, it will learn that if a coin is

of 3 g, it will be a 1 rupee coin. Let's

give a new coin to the machine. On the

basis of the weight of the new coin,

your model will predict the currency.

Hence, supervised learning uses labeled

data to train the model. Here, the

machine knew the features of the object

and also the labels associated with

those features. On this note, let's move

to unsupervised learning and see the

difference. Suppose you have cricket

data set of various players with their

respective scores and wickets taken.

When we feed this data set to the

machine, the machine identifies the

pattern of player performance. So, it

plots this data with the respective

wickets on the x-axis while runs on the

y-axis. While looking at the data,

you'll clearly see that there are two

clusters. The one cluster are the

players who scored high runs and took

less wickets while the other cluster is

of the players who scored less runs but

took many wickets. So here we interpret

these two clusters as batsmen and

bowlers. The important point to note

here is that there were no labels of

batsmen and bowlers. Hence the learning

with unlabeled data is unsupervised

learning. So we saw supervised learning

where the data was labeled and the

unsupervised learning where the data was

unlabeled. And then there is

reinforcement learning which is a

reward-based learning or we can say that

it works on the principle of feedback.

Here let's say you provide the system

with an image of a dog and ask it to

identify it. The system identifies it as

a cat. So you give a negative feedback

to the machine saying that it's a dog's

image. The machine will learn from the

feedback and finally if it comes across

any other image of a dog, it'll be able

to classify it correctly. That is

reinforcement learning. To generalize

machine learning model, let's see a

flowchart. Input is given to a machine

learning model which then gives the

output according to the algorithm

applied. If it's right, we take the

output as a final result. Else we

provide feedback to the training model

and ask it to predict until it learns. I

hope you've understood supervised and

unsupervised learning. So let's have a

quick quiz. You have to determine

whether the given scenarios uses

supervised or unsupervised learning.

Simple, right? Scenario one. Facebook

recognizes your friend in a picture from

an album of tagged photographs.

Scenario two, Netflix recommends new

movies based on someone's past movie

choices.

Scenario three, analyzing bank data for

suspicious transactions and flagging the

fraud transactions. Think wisely and

comment below your answers. Moving on,

don't you sometimes wonder how is

machine learning possible in today's

era? Well, that's because today we have

humongous data available. Everybody's

online either making a transaction or

just surfing the internet and that's

generating a huge amount of data every

minute. And that data my friend is the

key to analysis. Also, the memory

handling capabilities of computers have

largely increased which helps them to

process such huge amount of data at hand

without any delay. And yes, computers

now have great computational powers. So

there are a lot of applications of

machine learning out there. To name a

few, machine learning is used in

healthcare where diagnostics are

predicted for doctor's review. The

sentiment analysis that the tech giants

are doing on social media is another

interesting application of machine

learning. Fraud detection in the finance

sector and also to predict customer

churn in the e-commerce sector. While

booking a cab, you must have encountered

search pricing often where it says the

fair of your trip has been updated.

Continue booking. Yes, please. I'm

getting late for office. Well, that's an

interesting machine learning model which

is used by global taxi giant Uber and

others where they have differential

pricing in real time based on demand,

the number of cars available, bad

weather, rush hour, etc. So they use the

search pricing model to ensure that

those who need a cab can get one. Also,

it uses predictive modeling to predict

where the demand will be high with a

goal that drivers can take care of the

demand and search pricing can be

minimized. Great. Hey Siri, can you

remind me to book a cab at 6 p.m. today?

>> Okay, I'll remind you.

>> Thanks.

>> No problem.

>> Comment below some interesting everyday

examples around you where machines are

learning and doing amazing jobs. Hi

guys, this is Acha from SimplyLearn and

we're going to talk about the two types

of machine learning supervised and

unsupervised learning their types and

applications. But before we talk about

them, let's quickly understand what is

machine learning. These days

applications use artificial intelligence

in machine learning to optimize speech

recognition. I usually ask Siri things I

want to know like hey Siri how far is

the nearest subway? So whenever we ask

something to Siri, a powerful speech

recognition kicks off and converts the

audio into its corresponding textual

form which is then sent to the Apple

servers for further processing. Then

neural language processing algorithms

are run to understand the user's intent

and then finally Siri tells you the

answer. Well, this is what machine

learning is all about. making the

machines learn and act like humans by

feeding them with data and information

without being explicitly programmed. As

we saw in the previous example, when the

data comes in, machines immediately

starts analyzing the data and eventually

gets trained on it and learns it. Now

when a new data point comes in, machine

accurately makes prediction and

decisions based on the past data. Now

that you know what is machine learning,

let's talk about supervised and

unsupervised learning. Supervised

learning as the name suggests works

under supervision that is it's a

learning in which machine is trained

with data which is welllabeled and then

predicts with the help of the label data

set. But what is a label data set? Data

for which you already know the target

answer is called a label data. Like I

show you an image and tell you that it's

a dog then it's a label data. While if I

show you an image without telling you

what exactly it is, then it's an

unlabelled data. Now let's say we have

images which are labeled as spoon or

knife. We then feed it to the machine

which analyzes and learns the

association of these images with its

labels based on its features such as

shape, size, sharpness etc. Now when a

new image is fed to the machine without

any label with the help of the past data

the machine is able to predict

accurately and tell that it's a spoon.

Hence in supervised machine learning the

algorithm teaches the model to learn

from the labeled example that we

provide. So supervised learning can be

further divided into classification and

regression. It is a classification

problem when the output variable is

categoric such as red or blue, disease

or no disease, male or female. Whereas

it's a regression problem when the

output variable is a real or continuous

value. For example, salary based on work

experience, weight based on height. So

it creates a predictive model showing

trends in data. So now if I say will I

get a salary raise or not, that's

classification. But if I say how much

salary raise will I get, that is

regression. Now let's understand

classification with the help of an

example. In order to predict whether an

email is a spam or not, first we need to

teach our machine what a spam mail looks

like. This is done based on a lot of

spam filters like firstly reviewing the

content of the email then review the

email header and search if it contains

any falsified information. This is done

based on some keywords like free lottery

prize claim etc. Then general blacklist

filters to stop emails that come from

already blacklisted known spammers and

etc. So all these filters scores the

email which is known as spam score. The

lower the total spam score of the email,

it is more likely that the email will

land in subscribers inboxes. So based on

the content labels and spam score of the

new incoming mail, the algorithm decides

whether it should land in inbox or the

spam folder. Now let's quickly

understand regression. So let's say we

have two variables that is temperature

and humidity where temperature is the

independent variable and humidity is the

dependent variable such that as the

temperature increases humidity decreases

hence they're correlated. When we feed

this data to a regression model it will

understand the relationship between

these two variables and how one variable

depends on the other. After the machine

is trained it can easily predict the

humidity based on the given temperature.

Well, that was about regression. Now,

let's see some real life applications

where supervised learning is used. So,

supervised learning is used in risk

assessment to assess risk in financial

services or an insurance domain to

minimize the risk portfolio of the

companies. Image classification.

Facebook recognizes your friend in a

picture from an album of tact photos.

So, image classification is one of the

key use cases of demonstrating

supervised machine learning algorithms.

Obviously a lot more goes into all these

like convolutional neural networks etc.

Fraud detection whether the transactions

made by the user are authentic or not

and visual recognition the ability of a

machine learning model to identify

objects places people and actions in

images. Now let's quickly talk about

unsupervised learning. In unsupervised

learning, there is no supervision that

is no training will be given to the

machine allowing it to act on the data

which is not labeled. Hence, machine

tries to identify patterns and gives the

response. Let's take a similar example

as before. But this time we do not tell

the machine whether it's a spoon or a

knife. The machine identifies patterns

from the given set and groups them based

on their patterns, similarities, etc.

Again unsupervised learning can be

further grouped into clustering and

association. Clustering is basically

where the machine forms groups based on

the behavior of the data. Secondly,

association. It is a rule-based machine

learning to discover interesting

relation between variables in large data

sets. For example, which customer made

similar product purchases is clustering.

Whereas association is which products

were purchased together. Now let's

understand clustering with the help of

an example. To reduce their churn rate,

a telecom company studies the behavior

of the customers based on average call

duration and internet usage and observes

that while some customers call duration

is quite high, others have heavy

internet usage. The customers are

grouped based on their observed behavior

and a strategy is adopted to minimize

churn rate and maximize profit via

suitable promotions and campaigns. As

you can see in the chart on the right

hand side, customers in group A uses

more data and also have high

qualuration. Group B customers are heavy

internet users while group C customers

have high qualation. So group B will be

given more data benefit plans while

group C will be given cheaper call rates

to buy their loyalty. So this was the

example of clustering. Now let's

understand association with another

example. Let's say customer one goes to

a supermarket and buys these products.

say bread, milk, fruits, wheat. Then

customer two goes and buys bread, milk,

rice and butter. Now when customer three

goes and buys bread, it is highly likely

that he will also buy milk. Hence

relationship is established based on

customer behavior and recommendations

are made. Now let's look at some real

life applications of unsupervised

learning. Market basket analysis is a

machine learning model based on the

algorithm that if you buy a certain

group of items, you are less or more

likely to buy another group of items.

Semantic clustering. Semantically

similar words share similar context.

People post their queries on websites in

their own ways. Semantic clustering

groups all responses in a cluster with

same meaning to ensure that the customer

finds the information they want quickly

and easily. It plays an important role

in information retrieval. Good browsing

experience and comprehension. Delivery

store optimization. Machine learning

models are used to predict the demand

and keep up with the supply also to open

stores where demand is more and

optimizing routes for more efficient

deliveries according to past data and

behavior. We can also use unsupervised

machine learning models to identify

accidentprone areas based on the

intensity of those accidents and the

area in order to introduce safety

measures. By now I hope you've

understood supervised and unsupervised

learning. For a quick recap, let's see a

few differences between the two. The

most fundamental difference is that

supervised learning uses known and

labelled data and unsupervised learning

uses unlabelled data as their input.

Secondly, supervised learning follows a

feedback mechanism while unsupervised

learning does not. Also, the most

commonly used algorithms in supervised

learning are decision tree, logistic

regression, support vector machine, etc.

And in unsupervised learning there K

means clustering, hierarchical

clustering, a priori algorithm and many

more.

>> Welcome to linear regression. My name is

Richard Kersner. I'm with SimplyLearn.

Let's look at an example of a common use

for linear regression, profit estimation

of a company. If I was going to invest

in a company, I would like to know how

much money I could expect to make. So

we'll take a look at a venture

capitalist firm and try to understand

which companies they should invest in.

So we'll take the idea that we need to

decide the companies to invest in. We

need to predict the profit the company

makes and we're going to do it based on

the company's expenses and even just a

specific expense. In this case we have

our company, we have the different

expenses. So we have our R&D which is

your research and development. We have

our marketing. Uh we might have the

location. We might have what kind of

administration it's going through. Based

on all this different information, we

would like to calculate the profit. Now,

in actuality, there's usually about 23

to 27 different markers that they look

at if they're a heavy duty investor.

We're only going to take a look at one

basic one. We're going to come in and

for simplicity, let's consider a single

variable, R&D, and find out which

companies to invest in based on that.

So, we take our R&D and we're plotting

the profit based on the R&D expenditure,

how much money they put into the

research and development. And then we

look at the profit that goes with that.

We can predict a line to estimate the

profit. So we can draw a line right

through the data. And when you look at

that, you can see how much they invest

in the R&D is a good marker as to how

much profit they're going to have. We

can also note that companies spending

more on R&D make good profit. So let's

invest in the ones that spend a higher

rate in their R&D. What's in it for you?

First, we'll have an introduction to

machine learning followed by machine

learning algorithms. These will be

specific to linear regression and where

it fits into the larger model. Then

we'll take a look at applications of

linear regression, understanding linear

regression, and multiple linear

regression. Finally, we'll roll up our

sleeves and do a little programming in

use case profit estimation of companies.

Let's go ahead and jump in. Let's start

with our introduction to machine

learning along with some machine

learning algorithms and where that fits

in with linear regression. Let's look at

another example of machine learning.

Based on the amount of rainfall, how

much would be the crop yield? So we here

we have our crops, we have our rainfall,

and we want to know how much we're going

to get from our crops this year. So

we're going to introduce two variables,

independent and dependent. The

independent variable is a variable whose

value does not change by the effect of

other variables and is used to

manipulate the dependent variable. It is

often denoted as X. In our example,

rainfall is the independent variable.

This is a wonderful example because you

can easily see that we can't control the

rain, but the rain does control the

crop. So we talk about the independent

variable controlling the dependent

variable. Let's define dependent

variable as a variable whose value

change when there is any manipulation in

the values of the independent variables.

It is often denoted as y. And you can

see here our crop yield is dependent

variable and it is dependent on the

amount of rainfall received. Now that

we've taken a look at a real life

example, let's go a little bit into the

theory and some definitions on machine

learning and see how that fits together

with linear regression. numerical and

categorical values. Let's take our data

coming in and this is kind of random

data from any kind of project. We want

to divide it up into numerical and

categorical. So numerical is numbers,

age, salary, height, where categorical

would be a description, the color, a

dog's breed, gender. Categorical is

limited to very specific items where

numerical is a range of information. Now

that you've seen the difference between

numerical and categorical data, let's

take a look at some different machine

learning definitions. When we look at

our different machine learning

algorithms, we can divide them into

three areas. Supervised, unsupervised,

reinforcement. We're only going to look

at supervised today. Unsupervised means

we don't have the answers and we're just

grouping things. Reinforcement is where

we give positive and negative feedback

to our algorithm to program it. and it

doesn't have the information till after

the fact. But today we're just looking

at supervised because that's where

linear regression fits in. In supervised

data, we have our data already there and

our answers for a group. And then we use

that to program our model and come up

with an answer. The two most common uses

for that is through the regression and

classification. Now, we're doing linear

regression. So, we're just going to

focus on the regression side. And in the

regression we have simple linear

regression, we have multiple linear

regression and we have polomial linear

regression. Now on these three simple

linear regression is the examples we've

looked at so far where we have a lot of

data and we draw a straight line through

it. Multiple linear regression means we

have multiple variables. Remember where

we had the rainfall and the crops. We

might add additional variables in there

like how much food do we give our crops?

When do we harvest them? Those would be

additional information add into our

model and that's why it' be multiple

linear regression. And finally we have

polomial linear regression that is

instead of drawing a line we can draw a

curved line through it. Now that you see

where regression model fits into the

machine learning algorithms and we're

specifically looking at linear

regression. Let's go ahead and take a

look at applications for linear

regression. Let's look at a few

applications of linear regression.

Economic growth used to determine the

economic growth of a country or a state

in the coming quarter can also be used

to predict the GDP of a country. Product

price can be used to predict what would

be the price of a product in the future.

We can guess whether it's going to go up

or down or should I buy today. Housing

sales to estimate the number of houses a

builder would sell and what price in the

coming months. Score predictions.

Cricket fever to predict the number of

runs a player would score in the coming

matches based on the previous

performance. I'm sure you can figure out

other applications you could use linear

regression for. So let's jump in and

let's understand linear regression and

dig into the theory. Understanding

linear regression. Linear regression is

the statistical model used to predict

the relationship between independent and

dependent variables by examining two

factors. The first important one is

which variables in particular are

significant predictors of the outcome

variable. And the second one that we

need to look at closely is how

significant is the regression line to

make predictions with the highest

possible accuracy. If it's inaccurate,

we can't use it. So, it's very important

we find out the most accurate line we

can get. Since linear regression is

based on drawing a line through data,

we're going to jump back and take a look

at some uklitian geometry. The simplest

form of a simple linear regression

equation with one dependent and one

independent variable is represented by y

= m * x + c. And if you look at our

model here, we plotted two points on

here. Uh x1 and y1, x2 and y2. y being

the dependent variable, remember that

from before. And x being the independent

variable. So y depends on whatever x is.

M in this case is the slope of the line

where M equals the difference in the Y2

- Y1 and X2 - X1. And finally we have C

which is the coefficient of the line or

where happens to cross the zero axis.

Let's go back and look at an example we

used earlier of linear regression. We're

going to go back to plotting the amount

of crop yield based on the amount of

rainfall. And here we have our rainfall.

Remember, we cannot change rainfall. And

we have our crop yield, which is

dependent on the rainfall. So, we have

our independent and our dependent

variables. We're going to take this and

draw a line through it as best we can

through the middle of the data. And then

we look at that. We put the red point on

the y ais is the amount of crop yield

you can expect for the amount of

rainfall represented by the green dot.

So, if we have an idea what the rainfall

is for this year and what's going on,

then we can guess how good our crops are

going to be. and we've created a nice

line right through the middle to give us

a nice mathematical formula. Let's take

a look and see what the math looks like

behind this. Let's look at the intuition

behind the regression line. Now, before

we dive into the math and the formulas

that go behind this and what's going on

behind the scenes, I want you to note

that when we get into the case study and

we actually apply some Python script

that this math that you're going to see

here is already done automatically for

you. You don't have to have it

memorized. It is, however, good to have

an idea what's going on so if people

reference the different terms, you'll

know what they're talking about. Let's

consider a sample data set with five

rows and find out how to draw the

regression line. We're only going to do

five rows because if we did like the

rainfall with hundreds of points of

data, that would be very hard to see

what's going on with the mathematics.

So, we'll go ahead and create our own

two sets of data. And we have our

independent variable x and our dependent

variable y. And when x was 1, we got y =

2. When x was uh 2, y was 4. And so on

and so on. If we go ahead and plot this

data on a graph, we can see how it forms

a nice line through the middle. You can

see where it's kind of grouped going

upwards to the right. The next thing we

want to know is what the means is of

each of the data coming in, the x and

the y. The means doesn't mean anything

other than the average. So, we add up

all the numbers and divide by the total.

So, 1 + 2 + 3 + 4 + 5 over 5 equals 3.

And the same for y, we get four. If we

go ahead and plot the means on the

graph, we'll see we get 3a 4, which

draws a nice line down the middle, a

good estimate. Here, we're going to dig

deeper into the math behind the

regression line. Now, remember before I

said you don't have to have all these

formulas memorized or fully understand

them, even though we're going to go into

a little more detail of how it works.

And if you're not a math wiz and you

don't know if you've never seen the

sigma character before, which looks a

little bit like an e that's opened up,

that just means summation. That's all

that is. So, when you see the sigma

character, it just means we're adding

everything in that row. And for

computers, this is great because as a

programmer, you can easily iterate

through each of the XY points and create

all the information you need. So in the

top half, you can see where we've broken

that down into pieces. And as it goes

through the first two points, it

computes the squared value of X, the

squared value of Y, and X * Y. And then

it takes all of X and adds them up. All

of Y adds them up. All of X squar adds

them up. And so on and so on. And you

can see we have the sum of equal to 15.

The sum is equal to 20. all the way up

to x * y where the sum equals 66. This

all comes from our formula for

calculating a straight line where y

equals the slope* x plus the coefficient

c. So we go down below and we're going

to compute more like the averages of

these and we're going to explain exactly

what that is in just a minute and where

that information comes from. It's called

the square means error, but we'll go

into that in detail in a few minutes.

All you need to do is look at the

formula and see how we've gone about

computing it line by line instead of

trying to have a huge set of numbers

pushed into it. And down here you'll see

where the slope m equals and then the

top part if you read through the

brackets you have the number of data

points times the sum of x * y which we

computed one line at a time there. And

that's just the 66. and take all that

and you subtract it from the sum of x

times the sum of y and those have both

been computed. So you have 15 * 20. And

on the bottom we have the number of

lines times the sum of x^2 easily

computed as 86 for the sum minus I'll

take all that and subtract the sum of

x^2. And we end up as we come across

with our formula. You can plug in all

those numbers which is very easy to do

on the computer. You don't have to do

the math on a piece of paper or

calculator. And you'll get a slope of 6

and you'll get your C coefficient. If

you continue to follow through that

formula, you'll see it comes out as

equal to 2.2. Continuing deeper into

what's going behind the scenes, let's

find out the predicted values of y for

corresponding values of x using the

linear equation where m=6 and c = 2.2.

We're going to take these values and

we're going to go ahead and plot them.

We're going to predict them. So y =6 * x

= 1 + 2.2 = 2.8 so on and so on. And

here the blue points represent the

actual y values and the brown points

represent the predicted yv values based

on the model we created. The distance

between the actual and predicted values

is known as residuals or errors. The

best fit line should have the least sum

of squares of these errors also known as

equare. If we put these into a nice

chart where you can see X and you can

see Y what the actual values were and

you can see Y predicted you can easily

see where we take Y minus Y predicted

and we get an answer. What is the

difference between those two and if we

square that Y - Y prediction squared we

can then sum those squared values.

That's where we get the 64 plus the.36 +

1 all the way down until we have a

summation equals 2.4. So the sum of

squared errors for this regression line

is 2.4. We check this error for each

line and conclude the best fit line

having the least e value. In a nice

graphical representation, we can see

here where we keep moving this line

through the data points to make sure the

best fit line has the least squared

distance between the data points and the

regression line. Now we only looked at

the most commonly used formula for

minimizing the distance. There are lots

of ways to minimize the distance between

the line and the data points like sum of

squared errors, sum of absolute errors,

root mean square error, etc. What you

want to take away from this is whatever

formula is being used, you can easily

using a computer programming and

iterating through the data calculate the

different parts of it. That way, these

complicated formulas you see with the

different summations and absolute values

are easily computed one piece at a time.

Up until this point, we've only been

looking at two values, X and Y. Well, in

the real world, it's very rare that you

only have two values when you're

figuring out a solution. So, let's move

on to the next topic, multiple linear

regression. Let's take a brief look at

what happens when you have multiple

inputs. So, in multiple linear

regression, we have uh well, we'll start

with the simple linear regression where

we had y = m + x + c and we're trying to

find the value of y. Now with multiple

linear regression we have multiple

variables coming in. So instead of

having just x we have x1 x2 x3 and

instead of having just one slope each

variable has its own slope attached to

it. As you can see here we have m1 m2 m3

and we still just have the single

coefficient. So when you're dealing with

multiple linear regression you basically

take your single linear regression and

you spread it out. So you have y = m1 *

x1 + m2 * x2 so on all the way to m x to

the nth and then you add your

coefficient on there. Implementation of

linear regression. Now we get into my

favorite part. Let's understand how

multiple linear regression works by

implementing it in Python. If you

remember before we were looking at a

company and just based on its R&D trying

to figure out its profit. We're going to

start looking at the expenditure of the

company. We're going to go back to that.

We're going to predict his profit, but

instead of predicting it just on the

R&D, we're going to look at other

factors like administration costs,

marketing costs, and so on. And from

there, we're going to see if we can

figure out what the profit of that

company is going to be. To start our

coding, we're going to begin by

importing some basic libraries. And

we're going to be looking through the

data before we do any kind of linear

regression. We're going to take a look

at the data to see what we're playing

with. Then we'll go ahead and format the

data to the format we need to be able to

run it in the linear regression model.

And then from there we'll go ahead and

solve it and just see how valid our

solution is. So let's start with

importing the basic libraries. Now I'm

going to be doing this in Anaconda

Jupyter notebook, a very popular IDE. I

enjoy it because it's such a visual to

look at and so easy to use. Um just any

ID for Python will work just fine for

this. So break out your favorite Python

IDE. So, here we are in our Jupyter

notebook. Let me go ahead and paste our

first piece of code in there. And let's

walk through what libraries we're

importing. First, we're going to import

numpy as np. And then I want you to skip

one line and look at import pandas as

pd. These are very common tools that you

need with most of your linear

regression. The numpy, which stands for

number python, is usually denoted as np,

and you have to almost have that for

your sklearn toolbox. So, you always

import that right off the beginning.

pandas. Although you don't have to have

it for your sklearn libraries, it does

such a wonderful job of importing data,

setting it up into a data frame so we

can manipulate it rather easily and it

has a lot of tools also in addition to

that. So we usually like to use the

pandas when we can and I'll show you

what that looks like. The other three

lines are for us to get a visual of this

data and take a look at it. So we're

going to import mapplot library.pipplot

as plt and then seabour as sns. Seabor

works with the mattplot library. So you

have to always import mapplot library

and then seabour sits on top of it. And

we'll take a look at what that looks

like. You could use any of your own

plotting libraries you want. There's all

kinds of ways to look at the data. These

are just very common ones. And the

seabor is so easy to use. It just looks

beautiful. It's a nice representation

that you can actually take and show

somebody. And the final line is the

amberigned mattplot library inline. That

is only because I'm doing an inline IDE.

My interface in the Anaconda Jupiter

notebook requires I put that in there or

you're not going to see the graph when

it comes up. Let's go ahead and run

this. It's not going to be that

interesting because we're just setting

up variables. In fact, it's not going to

do anything that we can see, but it is

importing these different libraries and

setup. The next step is load the data

set and extract independent and

dependent variables. Now, here in the

slide, you'll see companies equals PD

read CSV. And it has a long line there

with the file at the end. 10,00

companies.csv. You're going to have to

change this to fit whatever setup you

have. And the file itself, you can

request. Just go down to the commentary

below this video and put a note in there

and SimplyLearn will try to get in

contact with you and supply you with

that file so you can try this coding

yourself. So, we're going to add this

code in here. And we're going to see

that I have companies equals

PD.reader_csv.

And I've changed this path to match my

computer. C/simplylearn/1000

companies.csv. And then below there,

we're going to set the x equals to

companies under the i location. And

because this is companies is a pd data

set, I can use this nice notation that

says take every row, that's what the

colon the first colon is, comma, except

for the last column. That's what the

second part is where we have a colon

minus one and we want the values set

into there. So x is no longer a data set

a pandas data set but we can easily

extract the data from our pandas data

set with this notation and then y we're

going to set equal to the last row. Well

the question is going to be what are we

actually looking at? So let's go ahead

and take a look at that and we're going

to look at the companies.

Which lists the first five rows of data

and I'll open up the file in just a

second so you can see where that's

coming from. But let's look at the data

in here as far as the way the pandas

sees it. When I hit run, you'll see it

breaks it out into a nice setup. This is

what pandas, one of the things pandas is

really good about is it looks just like

an Excel spreadsheet. You have your rows

and remember when we're programming, we

always start with zero. We don't start

with one. So it shows the first five

rows 0 1 2 3 4 and then it shows your

different columns. R&D spend,

administration, marketing spend, state,

profit. It even notes that the top are

column names. It was never told that,

but Pandas is able to recognize a lot of

things that they're not the same as the

data rows. Why don't we go ahead and

open this file up in a CSV so you can

actually see the raw data. So here I've

opened it up as a text editor. And you

can see at the top we have R&D spend,

administration, marketing spend, state,

profit, carriage return. I don't know

about you, but I'd go crazy trying to

read files like this. That's why we use

the pandas. You could also open this up

in an Excel and it would separate it

since it is a comma separated variable

file. But we don't want to look at this

one. We want to look at something we can

read rather easily. So let's flip back

and take a look at that top part, the

first five row. Now, as nice as this

format is where I can see the data, to

me it doesn't mean a whole lot. Maybe

you're an expert in business and

investments and you understand what

$165,34920

compared to the administration cost of

$136,897.80

so on so on helps to create the profit

of $192,26183.

That makes no sense to me whatsoever. No

pun intended. So let's flip back here

and take a look at our next set of code

where we're going to graph it so we can

get a better understanding of our data

and what it mean. So at this point we're

going to use a single line of code to

get a lot of information so we can see

where we're going with this. Let's go

ahead and paste that into our uh

notebook and see what we got going. And

so we have the visualization and again

we're using SNS which is pandas. As you

can see, we imported the mapplot

library.pipplot as plt, which then the

seabor uses and we imported the seabour

as sns. And then that final line of code

helps us show this in our um inline

coding. Without this, it wouldn't

display and you could display it to a

file and other means. And that's the map

plot library in line with the amber sign

at the beginning. So here we come down

to the single line of code. Seabor is

great because it actually recognizes the

panda data frame. So I can just take the

companies.core

for coordinates and I can put that right

into the seaborn. And when we run this,

we get this beautiful plot. And let's

just take a look at what this plot

means. If you look at this plot on mine,

the colors are probably a little bit

more purplish and blue than the original

one. Uh we have the columns and the

rows. We have R and D spending. We have

administration. We have marketing

spending and profit. And if you cross

index any two of these, since we're

interested in profit, if you cross-index

profit with profit, it's going to show

up, if you look at the scale on the

right, way up in the dark. Why? Because

those are the same data. They have an

exact correspondence. So R&D spending is

going to be the same as R&D spending.

And the same thing with administration

costs. But right down the middle, you

get this dark row or dark um diagonal

row that shows that this is the highest

corresponding data. That's exactly the

same. And as it becomes lighter, there's

less connections between the data. So we

can see with profit, obviously profit is

the same as profit. And next, it has a

very high correlation with R&D spending,

which we looked at earlier. And it has a

slightly less connection to marketing

spending and even less to how much money

we put into the administration. So now

that we have a nice look at the data,

let's go ahead and dig in and create

some actual useful linear regression

models so that we can predict values and

have a better profit. Now that we've

taken a look at the visualization of

this data, we're going to move on to the

next step. Instead of just having a

pretty picture, we need to generate some

hard data, some hard values. So let's

see what that looks like. We're going to

set up our linear regression model in

two steps. The first one is we need to

prepare some of our data so it fits

correctly. And let's go ahead and paste

this code into our Jupyter notebook. And

what we're bringing in is we're going to

bring in the sklearn pre-processing

where we're going to import the label

encoder and the one hot encoder. To use

the label encoder, we're going to create

a variable called label encoder and set

it equal to capital L label capital E

encoder. This creates a class that we

can reuse for transferring the labels

back and forth. Now about now you should

ask what labels are we talking about.

Let's go take a look at the data we

processed before and see what I'm

talking about here. If you remember when

we did the companies.head and we printed

the top five rows of data. We have our

columns going across. We have column

zero which is R&D spending, column one

which is administration, column two

which is marketing spending and column

three is state. And you'll see under

state we have New York, California,

Florida. Now to do a linear regression

model, it doesn't know how to process

New York. It knows how to process a

number. So the first thing we're going

to do is we're going to change that New

York, California, and Florida. And we're

going to change those to numbers. That's

what this line of code does here. X

equals and then it has the colon, 3 in

brackets. The first part, the colon,

comma, means that we're going to look at

all the different rows. So we're going

to keep them all together. But the only

row we're going to edit is the third

row. And in there, we're going to take

the label coder and we're going to fit

and transform the x also the third row.

So, we're going to take that third row,

we're going to set it equal to a

transformation. And that transformation

basically tells it that instead of

having a uh New York, it has a zero or a

one or a two. And then finally, we need

to do a one hot encoder, which equals

one hot encoder categorical features

equals three. And then we take the X and

we go ahead and do that equal to one hot

encoder fit transform X to array. This

final transformation preps our data for

us. So it's completely set the way we

need it as just a row of numbers. Even

though it's not in here, let's go ahead

and print X and just take a look what

this data is doing. You'll see I have an

array of arrays and then each array is a

row of numbers. And if I go ahead and

just do row zero, you'll see I have a

nice organized row of numbers that the

computer now understands. We'll go ahead

and take this out there because it

doesn't mean a whole lot to us. It's

just a row of numbers. Next on setting

up our data, we have avoiding dummy

variable trap. This is very important.

Why? Because the computer's

automatically transformed our header

into the setup and it's automatically

transformed all these different

variables. So when we did the encoder,

the encoder created two columns. And

what we need to do is just have the one

because it has both the variable and the

name. That's what this piece of code

does here. Let's go ahead and paste this

in here. And we have x= x colon, one

colon. All this is doing is removing

that one extra column we put in there

when we did our one hot encoder and our

label encoding. Let's go ahead and run

that. And now we get to create our

linear regression model. And let's see

what that looks like here. And we're

going to do that in two steps. The first

step is going to be in splitting the

data. Now, whenever we create a uh

predictive model of data, we always want

to split it up. So, we have a training

set and we have a testing set. That's

very important. Otherwise, we'd be very

unethical without testing it to see how

good our fit is. And then we'll go ahead

and create our multiple linear

regression model and train it and set it

up. Let's go ahead and paste this next

piece of code in here. And I'll go ahead

and shrink it down a size or two so it

all fits on one line. So from the

sklearn module selection, we're going to

import train test split. And you'll see

that we've created four completely

different variables. We have capital X

train capital X test smallercase Y train

smallerase Y test. That is the standard

way that they usually reference these

when we're doing different uh models.

usually see that a capital X and you see

the train and the test and the lowercase

Y. What this is is X is our data going

in. That's our R&D spin, our

administration, our marketing. And then

Y, which we're training, is the answer.

That's the profit because we want to

know the profit of an unknown entity. So

that's what we're going to shoot for in

this tutorial. The next part, train,

test, split. We take X and we take Y.

We've already created those. X has the

columns with the data in it and Y has a

column with profit in it. And then we're

going to set the test size equals 0.2.

That basically means 20%. So 20% of the

rows are going to be tested. We're going

to put them off to the side. So since

we're using a thousand lines of data,

that means that 200 of those lines we're

going to hold off to the side to test

for later. And then the random state

equals zero. We're going to randomize

which ones it picks to hold off to the

side. We'll go ahead and run this. It's

not overly exciting because it's setting

up our variables. But the next step is

the next step we actually create our

linear regression model. Now that we got

to the linear regression model, we get

that next piece of the puzzle. Let's go

ahead and put that code in there and

walk through it. So here we go. We're

going to paste it in there. And let's go

ahead and uh since this is a shorter

line of code, let's zoom up there so we

can get a good look. And we have from

the sklearn.linear_model,

we're going to import linear regression.

Now, I don't know if you recall from

earlier when we were doing all the math.

Let's go ahead and flip back there and

take a look at that. Do you remember

this where we had this long formula on

the bottom and we were doing all this

summization and then we also looked at

setting it up with the different lines

and then we also looked all the way down

to multiple linear regression where

we're adding all those formulas

together. All of that is wrapped up in

this one section. So what's going on

here is I'm going to create a variable

called regressor. And the regressor

equals the linear regression. That's a

linear regression model that has all

that math built in. So we don't have to

have it all memorized or have to compute

it individually. And then we do the

regressor.fit.

In this case, we do xrain and y train

because we're using the training data. X

being the data in and y being profit

what we're looking at. And this does all

that math for us. So within one click

and one line, we've created the whole

linear regression model and we fit the

data to the linear regression model. And

you can see that when I run the

regressor, it gives an output linear

regression. It says copy X equals true,

fit intercept equals true, in jobs equal

1, normalize equals false. It's just

giving you some general information on

what's going on with that regressor

model. Now that we've created our linear

regression model, let's go ahead and use

it. And if you remember, we kept a bunch

of data aside. So, we're going to do a Y

predict variable and we're going to put

in the X test. And let's see what that

looks like. Scroll up a little bit.

Paste that in here. Predicting the test

set results. So, here we have Y predict

equals regressor.predict

X test going in. And this gives us Y

predict. Now, because I'm in Jupiter in

line, I can just put the variable up

there. And when I hit the run button,

it'll print that array out. I could have

just as easily done print y predict. So

if you're in a different IDE that's not

an inline setup like the Jupyter

notebook, you can do it this way. Print

y predict. And you'll see that for the

200 different test variables we kept off

to the side, it's going to produce 200

answers. This is what it says the profit

are for those 200 predictions. But let's

don't stop there. Let's keep going and

take a couple look. We're going to take

just a short detail here and calculating

the coefficients and the intercepts.

This gives us a quick flash at what's

going on behind the line. We're going to

take a short detour here and we're going

to be calculating the coefficient and

intercepts. So you can see what those

look like. What's really nice about our

regressor we created is it already has

the coefficients for us. And we can

simply just print regressor.coefficient

underscore. When I run this, you'll see

our coefficients here. And if we can do

the regressor coefficient, we can also

do the regressor intercept. And let's

run that and take a look at that. This

all came from the multiple regression

model. And we'll flip over so you can

remember where this is going into where

it's coming from. You can see the

formula down here where y = m1 * x1 + m2

* x2 and so on and so on plus c the

coefficient. So these variables fit

right into this formula. Y equ= slope 1

* column 1 variable plus slope 2 *

column 2 variable all the way to the m

into the n and x to the n + c the

coefficient or in this case you have -

8.89 8 9 to the power of two etc etc

times the first column and the second

column and the third column and then our

intercept is the minus one3009

point. Boy, it gets kind of complicated

when you look at it. This is why we

don't do this by hand anymore. This is

why we have the computer to make these

calculations easy to understand and

calculate. Now, I told you that was a

short detour and we're coming towards

the end of our script. As you remember

from the beginning, I said if we're

going to divide this information, we

have to make sure it's a valid model,

that this model works and understand how

good it works. So calculating the R

squar value, that's what we're going to

use to predict how good our prediction

is. And let's take a look at what that

looks like in code. And so we're going

to use this from sklearn.metrics.

We're going to import R2 score. That's

the R squared value. We're looking at

the error. So in the R2 score, we take

our Y test versus our Y predict. Y test

is the actual values we're testing. That

was the one that was given to us. So we

know are true. The Y predict of those

200 values is what we think it was true.

And when we go ahead and run this, we

see we get a 9352.

That's the R2 score. Now, it's not

exactly a straight percentage. So it's

not saying it's 93% correct, but you do

want that in the upper 90s. O and higher

shows that this is a very valid

prediction based on the R2 score. And if

R squar value of N1 or 92 as we got on

our model remember it does have a random

generation involved. This proves the

model is a good model which means

success. Yay. We successfully trained

our model with certain predictors and

estimated the profit of the companies

using linear regression. What is

logistic regression? Let's say we have

to build a predictive model or a machine

learning model to predict whether the

passengers of the Titanic ship have

survived or not the shipwreck. So how do

we do that? So we use logistic

regression to build a model for this.

How do we use logistic regression? So we

have the information about the

passengers, their ID, whether they have

survived or not, their class and name

and so on and so forth. And we use this

information where we already know

whether the person has survived or not.

That is the labeled information and we

help the system to train based on this

information with based on this labeled

data. This is known as labeled data. And

during the process of building the

model, we probably will remove some of

the non-essential parameters or

attributes here. We only take those

attributes which are really required to

make these predictions. And once we

train the model, we run new data through

it whereby the model will predict

whether the passenger has survived or

not. All right. What is logistic

regression? As I mentioned earlier,

logistic regression is an algorithm for

performing binary classification. So

let's take an example and see how this

works. Let's say your car has not been

serviced for quite a few years and now

you want to find out if it is going to

break down in the near future. So this

is like a classification problem. Find

out whether your car will break down or

not. So how are we going to perform this

classification? So here's how it looks.

If we plot the information along the X

and Y axis, X is the number of years

since the last service was performed and

Y is the probability of your car

breaking down. And let's say this

information was this data rather was

collected from several car users. It's

not just your car but several car users.

So that is our labeled data. So the data

has been collected and um for for the

number of years and when the car broke

down and what was the probability and

that has been plotted along x and y

axis. So this provides an idea or from

this graph we can find out whether your

car will break down or not. We'll see

how. So first of all the probability can

go from 0 to one. As you all aware

probability can be between 0 and one.

And as we can imagine it is intuitive as

well. As the number of years are on the

lower side maybe 1 year, 2 years or 3

years till after the service the chances

of your car breaking down are very

limited. Right? So for example, chances

of your car breaking down or the

probability of your car breaking down

within 2 years of your last service are

0.1 probability. Similarly 3 years is

maybe.3 and so on. But as the number of

years increases let's say if it was 6 or

7 years there is almost a certainty that

your car is going to break down. That is

what this graph shows. So this is an

example of a application of the

classification algorithm and we will see

in little details how exactly logistic

regression is applied here. One more

thing needs to be added here is that the

dependent variables outcome is discrete.

So if we are talking about whether the

car is going to break down or not. So

that is a discrete value. The y that we

are talking about the dependent variable

that we are talking about what we are

looking at is whether the car is going

to break down or not yes or no that is

what we are talking about. So here the

outcome is discrete and not a continuous

value. So this is how the logistic

regression curve looks. Let me explain a

little bit what exactly and how exactly

we are going to uh determine the class

the outcome rather. So for a logistic

regression curve a threshold has to be

set saying that because this is a

probability calculation remember this is

a probability calculation and the

probability itself will not be zero or

one but based on the probability we need

to decide what the outcome should be. So

there has to be a threshold like for

example 0.5 can be the threshold let's

say in this case. So any value of the

probability below 0.5 is considered to

be zero and any value above.5 is

considered to be one. So an output of

let's say8

will mean that the car will break down.

So that is considered as an output of 1

and let's say an output of 29 is

considered as zero which means that the

car will not break down. So that's the

way logistic regression works. Now let's

do a quick comparison between logistic

regression and linear regression because

they both have the term regression in

them. So it can cause confusion. So

let's try to remove that confusion. So

what is linear regression? Linear

regression is a process is once again an

algorithm for supervised learning.

However, here you're going to find a

continuous value. You're going to

determine a continuous value. It could

be the price of a real estate property.

It could be your hike, how much hike

you're going to get or it could be a

stock price. These are all continuous

values. These are not discrete compared

to a yes or a no kind of a response that

we are looking for in logistic

regression. So this is one example of a

linear regression. Let's say the HR team

of a company tries to find out what

should be the salary hike of an

employee. So they collect all the

details of their existing employees,

their ratings and their salary hikes,

what has been given and that is the

labeled information that is available

and the system learns from this. It is

trained and it learns from this labeled

information so that when a new employees

information is fed based on the rating

it will determine what should be the

high. So this is a linear regression

problem and a linear regression example.

Now salary is a continuous value. You

can get 5,000, 5,500,

5,600. It is not discrete like a cat or

a dog or an apple or a banana. These are

discrete or a yes or a no. These are

discrete values, right? So this where

you're trying to find continuous values

is where we use linear regression. So

let's say just to extend on this

scenario, we now want to find out

whether this employee is going to get a

promotion or not. So we want to find out

that is a discrete problem, right? A yes

or no kind of a problem. In this case,

we actually cannot use linear regression

even though we may have labeled data. So

this is the label data. So based on the

employee rating these are the ratings

and then some people got the promotion

and this is the ratings for which people

did not get promotion that is a no and

this is the rating for which people got

promotion we just plotted the data about

whether a person has got an employee has

got promotion or not yes no right so

there is nothing in between and what is

the employees rating okay and ratings

can be continuous that is not an issue

but the output is discrete in In this

case whether employee got promotion yes

no okay so if we try to plot that and we

try to find a straight line this is how

it would look and as you can see it

doesn't look very right because looks

like there will be lot of errors this

root mean square error if you remember

for linear regression would be very very

high and also the the values cannot go

beyond zero or beyond one. So the graph

should probably look somewhat like this

clipped at 0 and one. But still the

straight line doesn't look right.

Therefore instead of using a linear

equation we need to come up with

something different and therefore the

logistic regression model looks somewhat

like this. So we calculate the

probability and if we plot that

probability not in the form of a

straight line but we need to use some

other equation. And we will see very

soon what that equation is. Then it is a

gradual process. Right? So you see here

people with some of these ratings are

not getting any promotions and then

slowly uh at certain rating they get

promotion. So that is a gradual process

and uh this is how the math behind

logistic regression looks. So we are

trying to find the odds for a particular

event happening and this is the formula

for finding the odds. So the probability

of an event happening divided by the

probability of the event not happening.

So P if it is the probability of the

event happening probability of the

person getting a promotion and divided

by the probability of the person not

getting a promotion that is 1 minus P.

So this is how you measure the odds. Now

the values of the odds range from 0 to

infinity. So when this probability is

zero then the odds will the value of the

odds is equal to zero and when the

probability becomes 1 then the value of

the odds is 1 by 0 that will be infinity

but the probability itself remains

between 0 and 1. Now this is how an

equation of a straight line looks. So y

is equal to beta 0 plus beta 1x where

beta 0 is the y intercept and beta 1 is

the slope of the line. If we take the

odds equation and take a log of both

sides, then this would look somewhat

like this. And the term logistic is

actually derived from the fact that we

are doing this. We take a log of px by 1

minus px. This is an extension of the

calculation of odds that we have seen,

right? And that is equal to beta 0 plus

beta 1x which is the equation of the

straight line. And now from here if you

want to find out the value of px we will

see we can take the exponential on both

sides and then if we solve that equation

we will get the equation of px like this

px is equal to 1 by 1 + e ^ of minus

beta 0 + beta 1x and recall this is

nothing but the equation of the line

which is equal to y is equal to beta 0 +

beta 1x. So that this is the equation

also known as the sigmoid function and

this is the equation of the logistic

regression alg. All right and if this is

plotted this is how the sigmoid curve is

obtained. So let's compare linear and

logistic regression how they are

different from each other. Let's go

back. So linear regression is solved or

used to solve regression problems and

logistic regression is used to solve

classification problems. So both are

called regression. But linear regression

is used for solving regression problems

where we predict continuous values.

Whereas logistic regression is used for

solving classification problems where we

have had to predict discrete values. The

response variables in case of linear

regression are continuous in nature.

Whereas here they are categorical or

discrete in nature. And the linear

regression helps to estimate the

dependent variable when there is a

change in the independent variable.

Whereas here in case of logistic

regression it helps to calculate the

probability or the possibility of a

particular event happening. And linear

regression as the name suggests is a

straight line. That's why it's called

linear regression. Whereas logistic

regression is a sigmoid function and the

curve is the shape of the curve is S.

It's an S-shaped curve. This is another

example of application of logistic

regression in weather prediction.

Whether it's going to rain or not rain.

Now keep in mind both are used in

weather prediction. If we want to find

the discrete values like whether it's

going to rain or not rain that is a

classification problem. We use logistic

regression. But if we want to determine

what is going to be the temperature

tomorrow, then we use linear regression.

So just keep in mind that in weather

prediction, we actually use both. But

these are some examples of logistic

regression. So we want to find out

whether it's going to be rain or not,

it's going to be sunny or not, whether

it's going to snow or not. These are all

logistic regression examples. A few more

examples. Classification of objects.

This is a again another example of

logistic regression. Now here of course

one distinction is that these are

multiclass classification. So logistic

regression is not used in its original

form but it is used in a slightly

different form. So we say whether it is

a dog or not a dog. I hope you

understand. So instead of saying is it a

dog or a cat or elephant we convert this

into saying so because we need to keep

it to binary classification. So we say

is it a dog or not a dog? Is it a cat or

not a cat? So that's the way logistic

regression can be used for classifying

objects. Otherwise there are other

techniques which can be used for

performing multiclass classification. In

healthcare logistic regression is used

to find the survival rate of a patient.

So they take multiple parameters like

trauma score and age and so on and so

forth and they try to predict the rate

of survival. All right. Now finally

let's take an example and see how we can

apply logistic regression to predict the

number that is shown in the image. So

this is actually a live demo. I will

take you into Jupyter notebook and u

show the code. But before that let me

take you through a couple of slides to

explain what we're trying to do. So

let's say you have an 8x8 image and the

the image has a number 1 2 3 4 and you

need to train your model to predict what

this number is. So how do we do this? So

the first thing is obviously in any

machine learning process you train your

model. So in this case we are using

logistic regression. So and then we

provide a training set to train the

model and then we test how accurate our

model is with the test data which means

that like any machine learning process.

We split our initial data into two parts

training set and test set. With the

training set we train our model and then

with the test set we we test the model

till we get good accuracy and then we

use it for for inference. Right? So that

is typical methodology of uh uh

training, testing and then deploying of

machine learning models. So let's uh

take a look at the code and uh see what

we are doing. So I'll not go line by

line but just take you through some of

the blocks. So first thing we do is

import all the libraries and then we

basically take a look at the images and

see what is the total number of images.

We can display using mattplot lip some

of the images or a sample of these

images and um then we split the data

into training and test as I mentioned

earlier and we can do some exploratory

analysis and uh then we build our model.

We train our model with the training set

and then we test it with our test set

and find out how accurate our model is

using the confusion matrix the heat map

and use heat map for visualizing this

and uh I will show you in the code what

exactly is the confusion matrix and how

it can be used for finding the accuracy

in our example we got we get an accuracy

of about 94 which is pretty good or 94%

which is pretty good all right so what

is the confusion matrix. This is an

example of a confusion matrix and uh

this is used for identifying the

accuracy of a classification model or

like a logistic regression model. So the

most important part in a confusion

matrix is that first of all this as you

can see this is a matrix and the size of

the matrix depends on how many outputs

uh we are expecting right. So the the

most important part here is that the

model will be most accurate when we have

the maximum numbers in its diagonal like

in this case that's why it has almost 93

94% because the diagonals should have

the maximum numbers and the others other

than diagonals the cells other than the

diagonals should have very few numbers.

So here that's what is happening. So

there is a two here. There are there's a

one here. But most of them are along the

diagonal. This what does this mean? This

means that the number that has been fed

is zero and the number that has been

detected is also zero. So the predicted

value and the actual value are the same.

So along the diagonals that is true.

Which means that let's let's take this

diagonal right. If if the maximum number

is here that means that uh like here in

this case it is 34 which means that 34

of the images that have been fed or

rather actually there are two

mclassifications in there. So 36 images

have been fed which have number four and

out of which 34 have been predicted

correctly as number four and one has

been predicted as number eight and

another one has been predicted as number

nine. So these are two mclassifications.

Okay. So that is the meaning of saying

that the maximum number should be in the

diagonal. So if you have all of them so

for an ideal model which has let's say

100% accuracy everything will be only in

the diagonal. There will be no numbers

other than zero in all other cells. So

that is like a 100% accurate model.

Okay. So that's the gist of how to use

this matrix. How to use this uh

confusion matrix. So I know the name uh

is a little funny sounding confusion

matrix but actually it is not very

confusing. It's very straightforward. So

you are just plotting what has been

predicted and what is the labeled

information or what is the actual data

that's also known as the ground truth

sometimes. Okay, these are some fancy

terms that are used. So predicted label

and the actual label that's all it is.

Okay. Yeah. So we are showing a little

bit more information here. So 38 have

been predicted and here you will see

that all of them have been predicted

correctly. There have been 38 zeros and

the predicted value and the actual value

is is exactly the same. Whereas in this

case right it has uh there are I think

37 + 5 yeah 42 have been fed the images

42 images are of digit three and uh the

accuracy is only 37 of them have been

accurately predicted. Three of them have

been predicted as number seven and two

of them have been predicted as number

eight and so on and so forth. Okay. All

right. So with that let's go into

Jupyter notebook and see how the code

looks. So this is the code in in Jupyter

notebook for logistic regression. In

this particular demo, what we are going

to do is train our model to recognize

digits which are the images which have

digits from let's say 0 to 5 or 0 to 9

and um and then we will see how well it

is trained and whether it is able to

predict these numbers correctly or not.

So let's get started. So the first part

is as usual we are importing some

libraries that are required and uh then

the last line in this block is to load

the digits. So let's go ahead and run

this code. Then here we will visualize

the shape of these uh digits. So we can

see here if we take a look this is how

the shape is 1797 by 64. These are like

8 by8 images. So that's that's what is

reflected in this uh shape. Now from

here onwards we are basically once again

importing some of the libraries that are

required like numpy and map plot and we

will take a look at uh some of the

sample images that we have loaded. So th

this one for example creates a figure uh

and then we go ahead and take a few

sample images to see how they look. So

let me run this code and so that it

becomes easy to understand. So these are

about five images sample images that we

are looking at 0 1 2 3 4. So this is how

the images this is how the data is.

Okay. And uh based on this we will

actually train our logistic regression

model and then we will test it and see

how well it is able to recognize. So the

way it works is the pixel information.

So as you can see here this is an 8x 8

pixel kind of a image and uh the each

pixel whether it is activated or not

activated that is the information

available for each pixel. Now based on

the pattern of this activation and

non-activation of the various pixels

this will be identified as a zero for

example right similarly as you can see

so overall each of these numbers

actually has a different pattern of the

pixel activation and that's pretty much

that our model needs to learn for which

number what is the pattern of the

activation of the pixels right so that

is what we are going to train our model.

Okay. So the first thing we need to do

is to split our data into training and

test data set. Right? So whenever we

perform any training, we split the data

into training and test. So that the

training data set is used to train the

system. So we pass this probably

multiple times. Uh and then we test it

with the test data set. And the split is

usually in the form of there and there

are various ways in which you can split

this data. It is up to the individual

preferences. In our case here we are

splitting in the form of 23 and 77. So

when we say test size as 2023

that means 23% of the entire data is

used for testing and the remaining 77%

is used for training. So there is a

readily available function which is uh

called train test split. So we don't

have to write any special code for the

splitting. It will automatically split

the data based on the proportion that we

give here which is test size. So we just

give the test size automatically

training size will be determined and uh

we pass the data that we want to split

and the the results will be stored in x

train and y train for the training data

set. And what is x train? This are these

are the features right which is like the

independent variable and y train is the

label right so in this case what happens

is we have the input value which is or

the features value which is in x train

and since this is a labeled data for

each of them each of the observations we

already have the label information

saying whether this digit is a zero or a

one or a two so that this this is what

will be used for comparison to find out

whether the the system is able to

recognize it correctly or there is an

error for each observation it will

compare with this right so this is the

label so the same way x train y train is

for the training data set x test y test

is for the test data set okay so let me

go ahead and execute this code as well

and then we can go and check quickly

what is the how many entries are there

and in each of this so x train the shape

is 1383x

64 and y train has 1383 because there is

uh nothing like the second part is not

required here and then x test shape we

see is 414 so actually there are 414

observations in test and 1383

observations in train so that's

basically what these four lines of code

are are saying okay then we import the

uh logistic regression

library and uh which is a part of

scikitlearn. So we we don't have to

implement the logistic regression

process itself. We just call these uh

the function and uh let me go ahead and

execute that so that uh we have the

logistic regression library imported.

Now we create an instance of logistic

regression. Right? So logistic regr is a

is an instance of logistic regression

and then we use that for training our

model. So let me first execute this

code. So these two lines. So the first

line basically creates an instance of

logistic regression model and then the

second line is where we are passing our

data the training data set. Right? This

is our the the predictors and uh this is

our target. We are passing this data set

to train our model. All right. So once

we do this in this case the data is not

large but by and large uh the training

is what takes usually a lot of time. So

we spend in machine learning activities

in machine learning projects we spend a

lot of time for the training part of it.

Okay. So here the data set is relatively

small so it was pretty quick. So all

right so now our model has been trained

using the training data set and uh we

want to see how accurate this is. So

what we'll do is we will test it out in

probably faces. So let me first try out

how well this is working for one image.

Okay, I will just try it out with one

image my the first entry in my test data

set and see whether it is uh correctly

predicting or not. So and in order to

test it so for training purpose we use

the fit method. There is a method called

fit which is for training the model and

once the training is done if you want to

test for uh a particular value new input

you use the predict method. Okay. So

let's run the predict method and we pass

this particular image and uh we see that

the shape is or the prediction is four.

So let's try a few more. Let me see for

the next 10 uh seems to be fine. So let

me just go ahead and test the entire

data set. Okay, that's basically what we

will do. So now we want to find out how

accurately this has u performed. So we

use the score method to find what is the

percentages of accuracy and we see here

that it has performed up to 94%

accurate. Okay. So that's uh on this

part. Now what we can also do is we can

um also see this accuracy using what is

known as confusion matrix. So let us go

ahead and uh try that as well. Uh so

that we can also visualize how well uh

this model has uh done. So let me

execute this piece of code which will

basically import some of the libraries

that are required and um we we basically

create a confusion matrix an instance of

confusion matrix by running confusion

matrix and passing these uh values. So

we have so this confusion_matrix

method takes two parameters one is the y

test and the other is uh the prediction.

So what is a y test? These are the

labeled values which we already know for

the test data set and predictions are

what the system has predicted for the

test data set. Okay. So this is known to

us and this is what the system has uh

the model has generated. So we kind of

create the confusion matrix and we will

print it. And uh this is how the

confusion matrix looks. As the name

suggests it is a matrix and um the key

point out here is that the accuracy of

the model is determined by how many

numbers are there in the diagonal. The

more the numbers in the diagonal, the

better the accuracy is. Okay. And first

of all, the total sum of all the numbers

in this whole matrix is equal to the

number of observations in the test data

set. That is the first thing, right? So

if you add up all these numbers, that

will be equal to the number of

observations in the test data set. And

then out of that, the maximum number of

them should be in the diagonal. That

means the accuracy is pretty good. If

the the numbers in the diagonal are less

and in all other places there are a lot

of numbers uh which means the accuracy

is very low. The diagonal indicates a

correct prediction that this means that

the actual value is same as the

predicted value. Here again actual value

is same as the predicted value and so

on. Right? So the moment you see a

number here that means the actual value

is something and the predicted value is

something else. Right? Similarly here

the actual value is something and the

predicted value is something else. So

that is basically how we read the

confusion matrix. Now how do we find the

accuracy? You can actually add up the

total values in the diagonal. So it it's

like 38 + 44 + 43 and so on and divide

that by the total number of test

observations that will give you the

percentage accuracy using a confusion

matrix. Now let us visualize this

confusion matrix in a slightly more

sophisticated way uh using a heat map.

So we will create a heat map with some

we'll add some colors as well. It's uh

it's like a more visually visually more

appealing. So that's the whole idea. So

if we let me run this piece of code and

this is how the heat map looks. Uh and

as you can see here the diagonals again

are all the values are here most of the

values. So which means reasonably this

seems to be reasonably accurate and yeah

basically the accuracy score is 94%.

This is calculated as I mentioned by

adding all these numbers divided by the

total test values or the total number of

observations in test data set. Okay. So

this is the confusion matrix for

logistic regression.

All right. So now that we have seen the

confusion matrix, let's take a quick

sample and see how well uh the system

has classified and we will take a a few

examples of the data. So if we see here

we we picked up randomly a few of them.

So this is uh number four which is the

actual value and also the predicted

value both are four. This is an image of

zero. So the predicted value is also

zero. Actual value is of course zero.

Then this is the image of nine. So this

has also been predicted correctly 9 and

actual value is 9. And this is the image

of one. And again this has been

predicted correctly as like the actual

value. Okay. So this was a quick demo of

logistic regression. How to use logistic

regression to identify images.

>> What is a decision tree? Let's go

through a very simple example before we

dig in deep. Decision tree is a

treeshaped diagram used to determine a

course of action. Each branch of the

tree represents a possible decision or

occurrence or reaction. Let's start with

a simple question. How to identify a

random vegetable from a shopping bag?

So, we have this group of vegetables in

here. And we can start off by asking a

simple question. Is it red? And if it's

not, then it's going to be the purple

fruit to the left, probably an eggplant.

If it's true, it's going to be one of

the red fruits. Is the diameter greater

than two? If false, it's going to be a

what looks to be a red chili. And if

it's true, it's going to be a bell

pepper from the capsicum family. So,

it's a capsicum.

Problems that decision tree can solve.

So, let's look at the two different

categories the decision tree can be used

on. It can be used on the

classification, the true false, yes, no,

and it can be used on regression where

we figure out what the next value is in

a series of numbers or a group of data.

In classification, the classification

tree will determine a set of logical if

then conditions to classify problems.

For example, discriminating between

three types of flowers based on certain

features. In regression, a regression

tree is used when the target variable is

numerical or continuous in nature. We

fit the regression model to the target

variable using each of the independent

variables. Each split is made based on

the sum of squared error. Before we dig

deeper into the mechanics of the

decision tree, let's take a look at the

advantages of using a decision tree and

we'll also take a glimpse at the

disadvantages. The first thing you'll

notice is that it's simple to

understand, interpret, and visualize. It

really shines here because you can see

exactly what's going on in a decision

tree. Little effort is required for data

preparation. So, you don't have to do

special scaling. There's a lot of things

you don't have to worry about when using

a decision tree. It can handle both

numerical and categorical data as we

discovered earlier and nonlinear

parameters don't affect its performance.

So even if the data doesn't fit an easy

curved graph, you can still use it to

create an effective decision or

prediction. If we're going to look at

the advantages of a decision tree, we

also need to understand the

disadvantages of a decision tree. The

first disadvantage is overfitting.

Overfitting occurs when the algorithm

captures noise in the data. That means

you're solving for one specific instance

instead of a general solution for all

the data. High variance. The model can

get unstable due to small variation in

data. Low bias tree. A highly

complicated decision tree tends to have

a low bias which makes it difficult for

the model to work with new data.

Decision tree important terms. Before we

dive in further, we need to look at some

basic terms. We need to have some

definitions to go with our decision tree

in the different parts we're going to be

using. We'll start with entropy. Entropy

is a measure of randomness or

unpredictability in the data set. For

example, we have a group of animals in

this picture. There's four different

kinds of animals. And this data set is

considered to have a high entropy. You

really can't pick out what kind of

animal it is based on looking at just

the four animals as a big clump of of uh

entities. So as we start splitting it

into subgroups, we come up with our

second definition which is information

gain. Information gain it is a measure

of decrease in entropy after the data

set is split. So in this case based on

the color yellow, we've split one group

of animals on one side as true and those

who aren't yellow as false. As we

continue down the yellow side, we split

based on the height. True or false

equals 10. And on the other side, height

is less than 10. True or false? And as

you see as we split it, the entropy

continues to be less and less and less.

And so our information gain is simply

the entropy E1 from the top and how it's

changed to E2 in the bottom. And we'll

look at the uh deeper math, although you

really don't need to know a huge amount

of math when you actually do the

programming in Python because it'll do

it for you. But we'll look on the actual

math of how they compute entropy.

Finally, we want to know the different

parts of our tree and they call the leaf

node. Leaf node carries the

classification or the decision. So it's

the final end at the bottom. The

decision node has two or more branches.

This is where we're breaking the group

up into different parts. And finally,

you have the root node. The topmost

decision node is known as the root node.

How does a decision tree work? Wonder

what kind of animals I'll get in the

jungle today? Maybe you're the hunter

with the gun. Or if you're more into

photography, you're a photographer with

a camera. So let's look at this group of

animals and let's try to classify

different types of animals based on

their features using a decision tree. So

the problem statement is to classify the

different types of animals based on

their features using a decision tree.

The data set is looking quite messy and

the entropy is high in this case. So

let's look at a training set or a

training data set and we're looking at

color. We're looking at height and then

we have our different animals. We have

our elephants, our giraffes, our

monkeys, and our tigers. And they're of

different colors and shapes. Let's see

what that looks like. And how do we

split the data? We have to frame the

conditions that split the data in such a

way that the information gain is the

highest. Note, gain is the measure of

decrease in entropy after splitting. So

the formula for entropy is the sum

that's what this symbol looks like. That

looks like kind of like a uh e funky e

of k where i equals 1 to k. K would

represent the number of animal the

different animals in there where value

or P value of I would be the percentage

of that animal times the log base 2 of

the same the percentage of that animal.

Let's try to calculate the entropy for

the current data set and take a look at

what that looks like. And don't be

afraid of the math. You don't really

have to memorize this math. Just be

aware that it's there and this is what's

going on in the background. And so we

have three giraffes, two tigers, one

monkey, two elephants, a total of eight

animals gathered. And if we plug that

into the formula, we get an entropy that

equals 3 over8. So we have three

giraffes, a total of 8 times the log.

Usually they use base 2 on the log. So

log base 2 of 3 over8 plus in this case,

let's say it's the elephants, 2 over 8.

Two elephants over total of 8 time log

base 2 2 over 8 plus one monkey over

total of 8. log base 2 1 over 8 and plus

2 over 8 of the tigers log base 2 over 8

and if we plug that into our computer or

calculator I obviously can't do logs in

my head we get an entropy equal to.571

the program will actually calculate the

entropy of the data set similarly after

every split to calculate the gain now

we're not going to go through each set

one at a time to see what those numbers

are just want you to be aware that this

is a formula or the mathematics behind

It gain can be calculated by finding the

difference of the subsequent entropy

values after a split. Now we will try to

choose a condition that gives us the

highest gain. We will do that by

splitting the data using each condition

and checking that the gain we get out of

them. The condition that gives us the

highest gain will be used to make the

first split. Can you guess what that

first split will be just by looking at

this image? As a human, it's probably

pretty easy to split it. Let's see if

you're right. If you guessed the color

yellow, you're correct. Let's say the

condition that gives us the maximum gain

is yellow. So we will split the data

based on the color yellow. If it's true,

that group of animals goes to the left.

If it's false, it goes to the right. The

entropy after the splitting has

decreased considerably. However, we

still need some splitting at both the

branches to attain an entropy value

equal to zero. So we decide to split

both the nodes using height as a

condition. Since every branch now

contains single label type, we can say

that entropy in this case has reached

the least value. And here you see we

have the giraffes, the tigers, the

monkey and the elephants all separated

into their own groups. This tree can now

predict all the classes of animals

present in the data set with 100%

accuracy. That was easy. Use case loan

repayment prediction. Let's get into my

favorite part and open up some Python

and see what the programming code and

the scripting looks like. In here, we're

going to want to do a prediction. And we

start with this individual here who's

requesting to find out how good his

customers are going to be, whether

they're going to repay their loan or not

for this bank. And from that, we want to

generate a problem statement to predict

if a customer will repay loan amount or

not. And then we're going to be using

the decision tree algorithm in Python.

Let's see what that looks like. And

let's dive into the code. In our first

few steps of implementation, we're going

to start by importing the necessary

packages that we need from Python. and

we're going to load up our data and take

a look at what the data looks like. So,

the first thing I need is I need

something to edit my Python and run it

in. So, let's flip on over. And here I'm

using the Anaconda Jupiter notebook.

Now, you can use any Python IDE you like

to run it in, but I find the Jupyter

Notebook's really nice for doing things

on the fly. And let's go ahead and just

paste that code in the beginning. And

before we start, let's talk a little bit

about what we're bringing in. And then

we're going to do a couple things in

here. where I have to make a couple

changes as we go through this first part

of the import. The first thing we bring

in is numpy as np. That's very standard

when we're dealing with mathematics,

especially with uh very complicated

machine learning tools. You'll almost

always see the numpy come in for your

num your numbers. It's called number

python. It has your mathematics in

there. In this case, we actually could

take it out, but generally you'll need

it for most of your different things you

work with. And then we're going to use

pandas as pd. That's also a standard.

The pandas is a dataf frame setup and

you can liken this to uh taking your

basic data and storing it in a way that

looks like an Excel spreadsheet. So as

we come back to this when you see np or

pd those are very standard uses you'll

know that that's the pandas and I'll

show you a little bit more when we

explore the data in just a minute. Then

we're going to need to split the data.

So I'm going to bring in our train test

and split and this is coming from the

sklearn package cross validation. In

just a minute, we're going to change

that and we'll go over that, too. And

then there's also the sktree import

decision tree classifier. That's the

actual tool we're using. Remember, I

told you don't be afraid of the

mathematics. It's going to be done for

you. Well, the decision tree classifier

has all that mathematics in there for

you, so you don't have to figure it back

out again. And then we have

sklearn.metrics

for accuracy score. We need to score our

our setup. That's the whole reason we're

splitting it between the training and

testing data. And finally, we still need

the sklearn import tree. And that's just

the basic tree function that's needed

for the decision tree classifier. And

finally, we're going to load our data

down here. And I'm going to run this and

we're going to get two things on here.

One, we're going to get an error. And

two, we're going to get a warning. Let's

see what that looks like. So the first

thing we had is we have an error. Why is

this error here? Well, it's looking at

this. It says I need to read a file. And

when this was written, the person who

wrote it, this is their path where they

stored the file. So let's go ahead and

fix that.

And I'm going to put in here my file

path. I'm just going to call it full

file name. And you'll see it's on my C

drive. And there's this very lengthy

setup on here where I stored the data

2.csv file.

Don't worry too much about the full path

because on your computer it'll be

different. The data.2 CSV file was

generated by SimplyLearn. If you want a

copy of that, you can comment down below

and request it here in the YouTube.

And then if I'm going to give it a name,

full file name, I'm going to go ahead

and change it here to full

file name. So let's go ahead and run it

now and see what happens.

And we get a warning

when you're coding. Understanding these

different warnings and these different

errors that come up is probably the

hardest lesson to learn. So let's just

go ahead and take a look at this and use

this as a uh opportunity to understand

what's going on here. If you read the

warning, it says the cross validation is

depreciated. So it's a warning on it's

being removed and it's going to be moved

in favor of the model selection. So if

we go up here, we have

sklearn.crossvalidation.

And if you research this and go to

sklearn site, you'll find out that you

can actually just swap it right in there

with model selection.

And so when I come in here and I run it

again, that removes a warning. What

they've done is they've had two

different developers develop it in two

different branches and then they decided

to keep one of those and eventually get

rid of the other one. That's all that is

and very easy and quick to fix.

Before we go any further, I went ahead

and opened up the data from this file.

Remember the the data file we just

loaded on here, the data_2.c

CSV. Let's talk a little bit more about

that and see what that looks like both

as a text file because it's a

commaepparated variable file and in a

spreadsheet. This is what it looks like

as a basic text file. You can see at the

top they've created a header and it's

got 1 2 3 4 five columns and each column

has data in it. And let me flip this

over cuz we're also going to look at

this uh in an actual spreadsheet so you

can see what that looks like. And here

I've opened it up in the open office

calc, which is pretty much the same as

um Excel and zoomed in. And you can see

we've got our columns and our rows of

data. A little easier to read in here.

We have a result, yes, yes, no. We have

initial payment, last payment, credit

score, house number. If we scroll way

down,

we'll see that this occupies a 101 lines

of code or lines of data with uh the

first one being a column and then 1,000

lines of data.

Now, as a programmer,

if you're looking at a small amount of

data, I usually start by pulling it up

in different sources so I can see what

I'm working with.

But in larger data, you won't have that

option. it would just be um too too

large. So you need to either bring in a

small amount that you can look at it

like we're doing right now or we can

start looking at it through the Python

code. So let's go ahead and move on and

take the next couple steps to explore

the data using Python. Let's go ahead

and see what it looks like in Python to

print the length and the shape of the

data. So let's start by printing the

length of the database. We can use a

simple lin function from Python. And

when I run this, you'll see that it's a

thousand long. And that's what we

expected. There's a thousand lines of

data in there. If you subtract the

column head, and this is one of the nice

things when we did the uh balance data

from the panda read CSV, you'll see that

the header is row zero. So, it

automatically removes a row and then

shows the data separate. It does a good

job sorting that data out for us. And

then we can use a different function.

And let's take a look at that. And

again, we're going to utilize the tools

in Panda.

And since the balance data was loaded as

a Panda data frame,

we can do a shape on it. And let's go

ahead and run the shape and see what

that looks like.

What's nice about the shape is not only

does it give me the length of the data,

we have a th00and lines, it also tells

me there's five columns. So when we were

looking at the data, we had five columns

of data. And then let's take one more

step to explore the data using Python.

And now that we've taken a look at the

length and the shape, let's go ahead and

use the uh pandas module for head.

Another beautiful thing in the data set

that we can utilize. So let's put that

on our sheet here. And we have print

data set and balance data.head.

And this is a pandas print statement of

its own. So it has its own print feature

in there. And then we went ahead and

gave a label for our print job here of

data set. Just a simple print statement.

And we run that. And let's just take a

closer look at that. Let me zoom in

here.

There we go.

Pandas does such a wonderful job of

making this a very clean readable data

set. So you can look at the data, you

can look at the column headers, you can

have it uh when you put it as the head,

it prints the first five lines of the

data. And we always start with zero. So

we have five lines. We have 0 1 2 3 4

instead of 1 2 3 4 5. That's a standard

scripting and programming set is you

want to start with the zero position.

And that is what the data head does. It

pulls the first five rows of data. Puts

it in a nice format that you can look at

and view. Very powerful tool to view the

data. So instead of having to flip and

open up an Excel spreadsheet or open

Office Cal or trying to look at a word

doc where it's all scrunched together

and hard to read, you can now get a nice

open view of what you're working with.

We're working with a shape of a thousand

long, five wide. So we have five columns

and we do the full data head. You can

actually see what this data looks like.

The initial payment, last payment,

credit scores, house number. So let's

take this now that we've explored the

data and let's start digging into the

decision tree. So in our next step,

we're going to train and build our data

tree. And to do that, we need to first

separate the data out. We're going to

separate into two groups so that we have

something to actually train the data

with. And then we have some data on the

side to test it to see how good our

model is. Remember with any of the

machine learning, you always want to

have some kind of test set to to weigh

it against so you know how good your

model is when you distribute it. Let's

go ahead and break this code down and

look at it in pieces. So first we have

our X and Y.

Where do X and Y come from? Well, X is

going to be our data and Y is going to

be the answer or the target. You can

look at it source and target. In this

case, we're using X and Y to denote the

data in and the data that we're actually

trying to guess what the answer is going

to be. And so to separate it, we can

simply put in X equals the balance of

the data values. The first brackets

means that we're going to select all the

lines in the database. So, it's all the

data. And the second one says we're only

going to look at columns 1 through five.

Remember, always start with zero. Zero

is a yes or no. And that's whether the

loan went default or not. So, we want to

start with one. If we go back up here,

that's the initial payment and it goes

all the way through the house number.

Well, if we want to look at uh 1 through

five, we can do the same thing for y,

which is the answers. And we're going to

set that just equal to the zero row. So,

it's just the zero row and then it's all

rows going in there. So, now we've

divided this into two different data

sets. One of them with the

data going in and one with the answers.

Next, we need to split the data.

And here you'll see that we have it

split into four different parts. The

first one is your X training, your X

test, your Y train, your Y test.

Simply put, we have X going in where

we're going to train it and we have to

know the answer to train it with. And

then we have X test where we're going to

test that data and we have to know in

the end what the Y was supposed to be.

And that's where this train test split

comes in that we loaded earlier in the

modules. This does it all for us. And

you can see they set the test size equal

to.3. So that's roughly 30% will be used

in the test. And then we use a random

state. So it's completely random which

rows it takes out of there. And then

finally we get to actually build our

decision tree. And they've called it

here CLF entropy. That's the actual

decision tree or decision tree

classifier. And in here, they've added a

couple variables which we'll explore in

just a minute. And then finally, we need

to fit the data to that. So, we take our

CLF entropy that we created and we fit

the X train. And since we know the

answers for X-ray or the Y train, we go

ahead and put those in. And let's go

ahead and run this. And what most of

these sklearn modules do is when you set

up the variable, in this case, when we

set the CLF entropy equal decision tree

classifier, it automatically prints out

what's in that decision tree. There's a

lot of variables you can play with in

here. And it's quite beyond the scope of

this tutorial to go through all of these

and how they work. But we're working on

entropy. That's one of the options.

We've added that it's completely a

random state of 100, so 100%. And we

have a max depth of three. Now, the max

depth, if you remember above when we

were doing the different graphs of

animals, means it's only going to go

down three layers before it stops. And

then we have minimal samples of leaves

is five. So, it's going to have at least

five leaves at the end. So, I'll have at

least three splits or have no more than

three layers and at least five end

leaves with the final result at the

bottom. Now that we've created our

decision tree classifier, not only

created it, but trained it, let's go

ahead and apply it and see what that

looks like. So, let's go ahead and make

a prediction and see what that looks

like. We're going to paste our predict

code in here. And before we run it,

let's just take a quick look at what's

doing here. We have a variable y predict

that we're going to do. And we're going

to use our variable CLF entropy that we

created.

And then you'll see predict. And it's

very common in the sklearn modules that

their different tools have the predict

when you're actually running a

prediction. In this case, we're going to

put our X test data in here. Now, if you

delivered this for use, an actual

commercial use, and distributed it, this

would be the new loans you're putting in

here to guess whether the person's going

to be uh pay them back or not. In this

case though, we need to test out the

data and just see how good our sample

is, how good of our tree does at

predicting the loan payments. And

finally, since Anaconda Jupyter notebook

is works as a command line for Python,

we can simply put the y predict en to

print it. I could just as easily have

put the print

and put brackets around y predict en to

print it out. We'll go ahead and do

that. It doesn't matter which way you do

it. And you'll see right here that it

runs a prediction. This is roughly 300

in here. Remember, it's 30% of a

thousand. So, you should have about 300

answers in here. And this tells you

which each one of those lines of ourh

test went in there. And this is what our

y predict came out. So, let's move on to

the next step where we're going to take

this data and try to figure out just how

good a model we have. So, here we go.

Since sklearn does all the heavy lifting

for you and all the math, we have a

simple line of code to let us know what

the accuracy is. And let's go ahead and

go through that and see what that means

and what that looks like. Let's go ahead

and paste this in. And let me zoom in a

little bit. There we go.

So you have a nice full picture. And

we'll see here. We're just going to do a

print accuracy is.

And then we do the accuracy score. And

this was something we imported um

earlier. If you remember at the very

beginning, let me just scroll up there

real quick so you can see where that's

coming from. That's coming from here

down here from sklearn.metrics metrics

import accuracy score. And you could

probably run a script, make your own

script to do this very easily. How

accurate is it? How many out of 300 do

we get right? And so we put in our y

test. That's the one we ran the predict

on. And then we put in our y predict en

that's the answers we got. And we're

just going to multiply that by 100

because this is just going to give us an

answer as a decimal and we want to see

it as a percentage. And let's run that

and see what it looks like. And if you

see here, we got an accuracy of

93.666667.

So when we look at the number of loans

and we look at how good our model fit,

we can tell people it has about a 93.6

fitting to it. So just a quick recap on

that. We now have accuracy set up on

here. And so we have created a model

that uses the decision tree algorithm to

predict whether a customer will repay

the loan or not. The accuracy of the

model is about 94.6%.

The bank can now use this model to

decide whether it should approve the

loan request from a particular customer

or not. And so this information is

really powerful. We may not be able to

as individuals understand all these

numbers because they have thousands of

numbers that come in, but you can see

that this is a smart decision for the

bank to use a tool like this to help

them to predict how good their uh

profit's going to be off of the loan

balances and how many are going to

default or not. We're going to be

looking at random forest, one of the

many powerful tools in the machine

learning library. Before we dive into

the topic, let's start by looking at a

few of the uses for random forest.

Currently today, it's used in remote

sensing. Uh for example, they're used in

the ETM devices. If you're a space buff,

that's the enhanced thermatic mapper

they use on satellites which see uh far

outside the human spectrum for looking

at land masses. and they acquire images

of the earth's surface. The accuracy is

higher and training time is less than

many other machine learning tools out

there. Also, object detection,

multiclass object detection is done

using random forest algorithms. A good

example is a traffic where you're trying

to sort out the different cars, buses,

and things. And it provides a better

detection in complicated environments.

They're very complicated up there. And

then we have uh another example connect.

And let's take a little closer look at

connect. Connect. They use a random

forest as part of the game console and

what it does is it tracks a body

movements and it recreates it in the

game and let's see what that looks like.

Uh we have a user who performs a step.

In this case it looks like Elvis Presley

going there that is then recorded so

that connect registers the movement and

then it marks the user based on

accuracy. And it looks like we have uh

Prince going on this one from Elvis

Presley to Prince. It's great. Uh so it

marks user base on the accuracy. If we

look at that a little closer, we have a

training set to identify body parts.

Where are the hands? Where are the feet?

Uh what's going on with the body? That

then goes into a random forest

classifier that learns from it. Once

we've trained the classifier, it then

identifies the body parts while the

person's dancing. It's able to represent

that in a computer format. And then

based on that, it scores the game and

how accurate you are as being Elvis

Presley or Prince in your dancing. So

why random forest? It's always important

to understand why we use this tool over

the other ones. What are the benefits

here? And so with the random forest, the

first one is there's no overfitting. If

you use of multiple trees, reduce the

risk of overfitting. Training time is

less. Overfitting means that we have fit

the data so close to what we have as our

sample that we pick up on all the weird

parts and instead of predicting the

overall data, you're predicting the

weird stuff which you don't want. High

accuracy runs efficiently on large

database. For large data, it produces

highly accurate predictions. In today's

world of uh big data, this is really

important. And this is probably where it

really shines. This is where Y random

forest really comes in. It estimates

missing data. Data in today's world is

very messy. So when you have a random

forest, it can maintain the accuracy

when a large proportion of the data is

missing. What that means is if you have

data that comes in from uh five or six

different areas and maybe they took one

set of statistics in one area and they

took a slightly different set of

statistics in the other. So they have

some of the sh same shared data, but one

is missing like the uh number of

children in the house if you're doing

something over demographics. and the

other one is missing the size of the

house. It will look at both of those

separately and build two different trees

and then it can do a very good job of

guessing which one fits better even

though it's missing that data. Let us

dig deep into the theory of exactly how

it works. And let's look at what is

random forest. Random forest or random

decision forest is a method that

operates by constructing multiple

decision trees. The decision of the

majority of the trees is chosen by the

random forest as the final decision. And

let's uh we have some nice graphics

here. We have a decision tree and they

actually use a real tree to denote the

decision tree which I love. And given a

random some kind of picture of a fruit.

This decision tree decides that the

output is it's an apple. And we have a

decision tree too where we have that

picture of the fruit goes in and this

one decides that it's a lemon. And the

decision three tree gets another image

and it decides it's an apple. And then

this all comes together in what they

call the random forest. And this random

forest then looks at it and says, "Okay,

I got two votes for apple, one vote for

lemon. The majority is apples. So the

final decision is apples." To understand

how the random forest works, we first

need to dig a little deeper and take a

look at the random forest and the actual

decision tree and how it builds that

decision tree. In looking closer at how

the individual decision trees work,

we'll go ahead and continue to use the

fruit example since we're talking about

trees and forests. A decision tree is a

treerehaped diagram used to determine a

course of action. Each branch of the

tree represents a possible decision,

occurrence, or reaction. So in here we

have a bowl of fruit and if you look at

that it looks like um they switch from

lemons to oranges. So we have oranges,

cherries, and apples. And the first

decision of the decision tree might be

is a diameter greater than or equal to

three. And if it says false, it knows

that they're cherries because everything

else is bigger than that. So all the

cherries fall into that decision. So we

have all that data we're training. We

can look at that. We know that that's

what's going to come up. Is the color

orange? Well, goes, hm, orange or red?

Well, if it's true, then it comes out as

the orange. And if it's false, that

leaves apples. So in this example, it

sorts out the fruit in the bowl or the

images of the fruit. A decision tree.

These are very important terms to know

because these are very central to

understanding the decision tree and when

working with them. The first is entropy.

Everything on the decision tree and how

it makes a decision is based on entropy.

Entropy is a measure of randomness or

unpredictability in the data set. uh

then they also have information gain,

the leaf node, the decision node and the

root node. We'll cover these other four

terms as we go down the tree, but let's

start with entropy. So starting with

entropy, we have here a high amount of

randomness. What that means is that

whatever is coming out of this decision,

if it was going to guess based on this

data, it wouldn't be able to tell you

whether it's a lemon or an apple. it

would just say it's a fruit. Uh so the

first thing we want to do is we want to

split this apart and we take the initial

data set. We're going to set create a

data set one and a data set two. We just

split it in two. And if you look at

these new data sets after splitting

them, the entropy of each of those sets

is much less. So for the first one,

whatever comes in there, it's going to

sort that data and it's going to say,

okay, if this data goes this direction,

it's probably an apple. And if it goes

into the other direction, it's probably

a lemon. So that brings us up to

information gain. It is the measure of

decrease in the entropy after the data

set is split. What that means in here is

that we've gone from one set which has a

very high entropy to two lower sets of

entropy and we've added in the values of

E1 for the first one and E2 for the

second two which are much lower. And so

that information gain is increased

greatly in this example. And so you can

find that the information grain simply

equals uh decision E1 minus E2. As we're

going down our list of uh definitions,

we'll look at the leaf node. And the

leaf node carries the classification or

the decision. So we look down here to

the leaf node. We finally get to our set

one or our set two. When it comes down

there and it says, "Okay, this object's

gone into set one." If it's gone into

set one, it's going to be split by some

means and we'll either end up with

apples on the leaf node or a lemon on

the leaf node. And on the right, it'll

either be an apple or lemons. Those leaf

nodes are those final decisions or

classifications. Uh that's the

definition of leaf node in here. If

we're going to have a final leaf where

we make the decision, we should have a

name for the nodes above it. And they

call those decision nodes. A decision

node. decision node has two or more

branches and you can see here where we

have the uh five apples and one lemon

and in the other case the five lemons

and one apple. They have to make a

choice of which tree it goes down based

on some kind of measurement or

information given to the tree. And that

brings us to our last definition. The

root node, the topmost decision node is

known as the root node. And this is

where you have all of your data and you

have your first decision. it has to make

or the first split in information. So

far, we've looked at a very general

image um with the fruit being split.

Let's look and see exactly what that

means to split the data and how do we

make those decisions on there. Uh let's

go in there and find out how does a

decision tree work. So let's try to

understand this and let's use a simple

example and we'll stay with the fruit.

We have a bowl of fruit and so let's

create a problem statement and the

problem is we want to classify the

different types of fruits in the bowl

based on different features. The data

set in the bowl is looking quite messy

and the entropy is high in this case. So

if this bowl was our decision maker, it

wouldn't know what choice to make. It

has so many choices. Which one do you

pick? Apple, grapes, or lemons. And so

we look in here. We're going to start

with a d a training set. So this is our

data that we're training our data with

and we have a number of options here. We

have the color and under the color we

have red yellow purple uh we have a

diameter uh 331 331 and we have a label

apple lemon grapes apple lemon grapes

and how do we split the data? We have to

frame the conditions to split the data

in such a way that the information gain

is the highest. It's very key to note

that we're looking for the best gain. We

don't want to just start sorting out the

smallest piece in there. We want to

split it the biggest way we can. And so

we measure this decrease in entropy.

That's what they call it, entropy.

There's our entropy after splitting. And

now we'll try to choose a condition that

gives us the highest gain. We will do

that by splitting the data using each

condition and checking the gain that we

get out of them. The conditions that

give us the highest gain will be used to

make the first split. So let's take a

look at these different conditions. We

have color, we have diameter, and if we

look underneath that, we have a couple

different values. is we have diameter

equals 3, color equals yellow, red,

diameter equals 1. And when we look at

that, you'll see over here we have 1 2 3

4 threes. That's a pretty hardy

selection. So let's say the condition

gives us a maximum gain of three. So we

have the most pieces fall into that

range. So our first split from our

decision node is we split the data based

on the diameter. Is it greater than or

equal to three? If it's not, that's

false. It goes into the grape bowl. And

if it's true, it goes into a bowl fold

of lemon and apples. The entropy after

splitting has decreased considerably. So

now we can make two decisions. If you

look at they're very uh much less chaos

going on there. This node has already

attain an entropy value of zero. As you

can see, there's only one kind of label

left for this branch. So no further

splitting is required for this node.

However, this node on the right is still

requires a split to decrease the entropy

further. So, we split the right node

further based on color. If you look at

this, if I split it on color, that

pretty much cuts it right down the

middle. And it's the only thing we have

left in our choices of color and

diameter, too. And if the color is

yellow, it's going to go to the right

bowl. And if it's false, it's going to

go to the left bowl. So, the entropy in

this case is now zero. So, now we have

three bowls with zero entropy. There's

only one type of data in each one of

those bowls. So, we can predict a lemon

with 100% accuracy. And we can predict

the apple also with 100% accuracy along

with our grapes up there. So, we've

looked at kind of a basic tree in our

forest. But what we really want to know

is how does a random forest work as a

whole. So to begin our um random forest

classifier, let's say we already have

built three trees. And we're going to

start with the first tree that looks

like this. Just like we did in the

example, this tree looks at the

diameter. If it's greater than or equal

to three, it's true. Otherwise, it's

false. So one side goes to the smaller

diameter, one side goes to larger

diameter. And if the color is orange,

it's going to go to the right. True.

We're using oranges now instead of

lemons. And if it's red, it's going to

go to the left. False. We build a second

tree very similar, but it's split

differently. Instead of the first one

being split by a diameter, uh this one

when they created it, if you look at

that first bowl, it has a lot of red

objects. So it says, is the color red?

Because that's going to bring our

entropy down the fastest. And so, of

course, if it's true, it goes to the

left. If it's false, it goes to the

right. And then it looks at the shape,

false or true, and so on and so on. And

tree three is the diameter equal to one.

And it came up with this because there's

a lot of cherries in this bowl. So that

would be the biggest split on there is

is the diameter equal to one. That's

going to drop the entropy the quickest.

And as you can see, it splits it into

true. If it goes false, and they've

added another category, does it grow in

the summer? And if it's false, it goes

off to the left. If it's true, it goes

off to the right. Let's go ahead and

bring these three trees so you can see

them all in one image. So this would be

three completely different trees

categorizing a fruit. And let's take a

fruit. Now let's try this. And this

fruit, if you look at it, we've

blackened it out. You can't see the

color on it. So it's missing data.

Remember one of the things we talked

about earlier is that a random forest

works really good if you're missing

data, if you're missing pieces. So this

fruit has an image, but maybe it's a

person had a black and white camera when

they took the picture. And we're going

to take a look at this. And it's going

to have um they put the color in there,

so ignore the color down there. But the

diameter equals three. We find out it

grows in the summer equals yes. And the

shape is a circle. And if you go to the

right, you can look at what one of the

decision trees did. This is the third

one. Is the diameter greater than equal

to three? Is a color orange? Well, it

doesn't really know on this one, but it

if you look at the value, it' say true,

and it go to the right. Tree two

classifies it as cherries. Is a color

equal red? Is the shape a circle? True.

It is a circle. So, this would look at

it and say, "Oh, that's a cherry." And

then we go to the other classifier and

it says, "Is the diameter equal one?"

Well, that's false. Does it grow in the

summer? True. So, it goes down and looks

at as oranges. So, how does this random

forest work? The first one says it's an

orange. The second one said it was a

cherry. And the third one says, hm, it's

an orange. And you can guess that if you

have two oranges and one says it's a

cherry, uh, when you add that all

together, the majority of the vote says

orange. So, the answer is it's

classified as an orange, even though we

didn't know the color and we're missing

data on it. I don't know about you, but

I'm getting tired of fruit. So, let's

switch. And I did promise you we'd start

looking at a case example and get into

some Python coding. Today, we're going

to use the case the iris flower

analysis.

This is the exciting part as we roll up

our sleeves and actually look at some

Python coding. Before we start the

Python coding, we need to go ahead and

create a problem statement. Wonder what

species of iris do these flowers belong

to? Let's try to predict the species of

the flowers using machine learning in

Python. Let's see how it can be done. So

here we begin to go ahead and implement

our Python code. And you'll find that

the first half of our implementation is

all about organizing and exploring the

data coming in. Let's go ahead and take

this first step, which is loading the

different modules into Python. And let's

go ahead and put that in our favorite

editor, whatever your favorite editor

is. In this case, I'm going to be using

the Anaconda Jupiter Notebook, which is

one of my favorites. Certainly, there's

Notepad++ and Eclipse and dozens of

others, or just even using the Python

terminal window. any of those will work

just fine to go ahead and explore this

Python coding. So, here we go. Let's go

ahead and flip over to our Jupyter

notebook. And I've already opened up a

new page for Python 3 code. And I'm just

going to paste this right in there. And

let's take a look and see what we're

bringing into our Python. The first

thing we're going to do is from the

sklearn.data sets import load iris. Now,

this isn't the actual data. So this is

just the module that allows us to bring

in the data, the load iris. And the iris

is so popular. It's been around since

1936 when Ronald Fiser published a paper

on it. And they're measuring the

different parts of the flower. And based

on those measurements, predicting what

kind of flower it is. And then if we're

going to do a random forest classifier,

we need to go ahead and import a random

forest classifier from the sklearn

module. So sklearn.semble

import random forest classifier. And

then we want to bring in two more

modules. Um, and these are probably the

most commonly used modules in Python and

data science with any of the um, other

modules that we bring in. And one is

going to be pandas. We're going to

import pandas as pd. PD is the common

term used for pandas. And pandas is

basically creates a data format for us

where when you create a pandas data

frame, it looks like an Excel

spreadsheet. And you'll see that in a

minute when we start digging deeper into

the code. Panda is just wonderful

because it plays nice with all the other

modules in there. And then we have

Numpy, which is our numbers Python. And

the numbers Python allows us to do

different mathematical sets on here.

We'll see right off the bat, we're going

to take our NP and we're going to go

ahead and seed the randomness with it

with zero. So NP.random seed is seeding

that as zero. This code doesn't actually

show anything. We're going to go ahead

and run it because I need to make sure I

have all those loaded. And then let's

take a look at the next module on here.

The next six slides, including this one,

are all about exploring the data.

Remember, I told you half of this is

about looking at the data and getting it

all set. So, let's go ahead and take

this code right here, the script, and

let's get that over into our Jupyter

notebook. And here we go. We've gone

ahead and uh run the imports. Now I'm

going to paste the code down here

and let's take a look and see what's

going on. The first thing we're doing is

we're actually loading the iris data.

And if you remember up here, we loaded

the module that tells it how to get the

iris data. Now we're actually assigning

that data to the variable iris. And then

we're going to go ahead and use the df

to define dataf frame. And that's going

to equal pd. And if you remember that's

pandas as pd. So that's our pandas and

panda dataf frame. And then we're

looking at iris data and columns equals

iris feature names. And we're going to

do the DF head. And let's run this so

you can understand what's going on here.

The first thing you want to notice is

that our DF has created uh what looks

like an Excel spreadsheet. And in this

Excel spreadsheet, we have set the

columns. So up on the top, you can see

the four different columns. And then we

have the data iris.data down below. It's

a little confusing without knowing where

this data is coming from. So let's look

at the bigger picture and I'm going to

go print. I'm just going to change this

for a moment and we're going to print

all of Iris and see what that looks

like. So when I print all of Iris I get

this long list of information. And you

can scroll through here and see all the

different titles on there. What's

important to notice is that first off

there's a brackets at the beginning. So

this is a Python dictionary

and in a Python dictionary you'll have a

key or a label and this label pulls up

whatever information comes after it. So

feature names which we actually used

over here under columns is equal to an

array of sele length sele width pedal

length pedal width. These are the

different names they have for the four

different columns. And if you scroll

down far enough you'll also see data

down here. Oh goodness, it came up right

towards the top. And uh data is equal to

the different data we're looking at.

Now, there's a lot of other things in

here like target. We're going to be

pulling that up in a minute. And there's

also the names uh the target names which

is further down. And we'll show you that

also in a minute. Let's go ahead and set

that back to the head. And this is one

of the neat features of pandas and panda

dataf frames is when you do df.ad or the

panda dataf frame. head. It'll print the

first five lines of the data set in

there along with the headers if you have

them. In this case, we have the column

headers set to iris features. And in

here, you'll see that we have 0 1 2 3 4.

In Python, most arrays always start at

zero. So, when you look at the first

five, it's going to be 0 1 2 3 4, not 1

2 3 4 5. So, now we've got our iris data

imported into a data frame. Let's take a

look at the next piece of code in here.

And so in this section here of the code,

we're going to take a look at the

target. And let's go ahead and get this

into our notebook, this piece of code,

so we can discuss it a little bit more

in detail. So here we are in our Jupyter

notebook. I'm going to put the code in

here. And before I run it, I want to

look at a couple things going on. So we

have uh DF species. And this is

interesting because right here you'll

see where I have DF species in brackets

which is uh the key code for creating

another column. And here we have

iris.target.

Now these are both in the pandas setup

on here. So in pandas we can do either

one. I could have just as easily done

iris and then in brackets target

depending on what I'm working on. Both

are um acceptable. Let's go ahead and

run this code and see how this changes.

And what we've done is we've added the

target from the iris data set as another

column on the end.

Now what species is this is what we're

trying to predict. So we have our data

which tells us the answer for all these

different pieces. And then we've added a

column with the answer. So that way when

we do our final setup, we'll have the

ability to program our our neural

network to look for these this different

data and know what a satossa is or a

veraricolor which we'll see in just a

minute or virginica. Those are the three

that are in there. And now we're going

to add one more column. I know we're

organizing all this data over and over

again. It's kind of fun. There's a lot

of ways to organize it. What's nice

about putting everything onto one data

frame is I can then do a print out and

it shows me exactly what I'm looking at.

And I'll show you where you where that's

different where you can alter that and

do it slightly differently. But let's go

ahead and put this into our script up to

now. And here we go. We're going to put

that down here and we're going to run

that. And let's talk a little bit about

what we're doing. Now we're exploring

data. And one of the challenges is

knowing how good your model is. Did your

model work? And to do this, we need to

split the data. And we split it into two

different parts. They usually call it

the training and the testing. And so in

here, we're going to go ahead and put

that in our database so you can see it

clearly. And we've set it df. And

remember, you can put brackets. This is

creating another column. Is train. So

we're going to use part of it for

training. And this equals np. Remember

that stands for numpy.random.uniform.

So we're generating a random number

between zero and one. And we're going to

do it for each of the rows. That's where

the length df comes from. So each row

gets a generated number. And if it's

less than 75, it's true. And if it's

greater than 75, it's false. This means

we're going to take 75% of the data

roughly because there's a randomness

involved. And we're going to use that to

train it. And then the other 25% we're

going to hold off to the side and use

that to test it later on. So let's flip

back on over and see what the next step

is. So now that we've labeled our

database for which is training and which

is testing, let's go ahead and sort that

into two different variables, train and

test. And let's take this code and let's

bring it into our project. And here we

go. Let's paste it on down here. And

before I run this, let's just take a

quick look at what's going on here. is

we have up above we created remember

there's our def head which prints the

first five rows and we've added a column

is train at the end and so we're going

to take that we're going to create two

variables we're going to create two new

data frames one's called train one's

called test 75% in train 25% in test and

then to sort that out we're going to do

that by doing df our main original data

frame with the iris data in it and if df

F is train equals true, that's going to

go in the train. And if DF is train

equals false, it goes in the test. And

so when I run this, we're going to print

out the number in each one. Let's see

what that looks like. And you'll see

that it puts 118 in the training module

and it puts 32 in the testing module,

which lets us know that there was 150

lines of data in here. So if you went

and looked at the original data, you

could see that there's 150 lines and

that's roughly 75% in one and 25% for us

to test our model on afterward. So let's

jump back to our code and see where this

goes. In the next two steps, we want to

do one more thing with our data, and

that's make it readable to humans. Um, I

don't know about you, but I hate looking

at zeros and ones. So, let's start with

the features and let's go ahead and take

those and make those readable to humans

and let's put that in our code.

Let's see. Here we go. Paste it in. And

you'll see here we've done a couple very

basic things. We know that the columns

in our data frame, again, this is a

panda thing, the DF columns, and we know

the first four of them, 01, 2, 3, that'd

be the first four are going to be the

features or the titles of those columns.

And so when I run this, you'll see down

here that it creates an index, sea

length, sea width, pedal length, and

pedal width. And this should be familiar

because if you look up here, here's our

column titles going across. And here's

the first four. One thing I want you to

notice here is that when you're in a

command line, whether it's Jupyter

notebook or you're running command line

in the uh terminal window, if you just

put the name of it, it'll print it out.

This is the same as doing print

features.

And the shortand is you just put

features in here. If you're actually

writing a code and saving the script and

running it by remote, you really need to

put the print in there. But for this,

when I run it, you'll see it gives me

the same thing.

But for this, we want to go ahead and

we'll just leave it as features because

it doesn't really matter. And this is

one of the fun thing about Jupyter

Notebooks is I'm just building the code

as we go. And then we need to go ahead

and create the labels for the other

part. So, let's take a look and see what

that for. Our final step in prepping our

data before we actually start running

the training and the testing is we're

going to go ahead and convert the

species on here into something the

computer understands. So, let's put this

code into our script and see where that

takes us.

All right, here we go. We've set y equal

to pd.factorize

train species of zero. So, let's break

this down just a little bit. We have our

pandas right here. PD factoriize. What

is factorized doing? I'm going to come

back to that in just a second. Let's

look at what train species is and why

we're looking at the group zero on

there. And let's go up here. And here is

our species.

Remember this on that? We created this

whole column here for species. And then

it has satossa, satossa, satossa,

satossa. And if you scroll down enough,

you'd also see virginica and

veraricolor.

We need to convert that into something

the computer understands. Zeros and

ones. So the train species of zero

because this is in the format of a of an

array of arrays. So you have to have the

zero on the end. And then species is

just that column. Factoriize goes in

there and looks at the fact that there's

only three of them. So when I run this,

you'll see that Y generates an array

that's equal to, in this case, it's the

training set, and it's zeros, ones, and

twos representing the three different

kinds of flowers we have. So now we have

something the computer understands, and

we have a nice table that we can read

and understand. And now finally we get

to actually start doing the predicting.

So here we go. Uh we have two lines of

code. Oh my goodness, that was a lot of

work to get to two lines of code. But

there is a lot in these two lines of

code. So let's take a look and see

what's going on here and put this into

our full script that we're running. And

let's paste this in here. And let's take

a look and see what this is. We have

we're creating a variable CLF. And we're

going to set this equal to the random

forest classifier. And we're passing two

variables in here. And there's a lot of

variables you can play with. As far as

these two are concerned, they're very

standard. In jobs, all that does is to

prioritize it. Not something to really

worry about. Usually when you're doing

this on your own computer, you do end

jobs equals 2. If you're working in a

larger or big data and you need to

prioritize it differently, this is what

that number does is it changes your

priorities and how it's going to run

across the system and things like that.

And then the random state is just how it

starts. Zero is fine for here.

But uh let's go ahead and run this.

We also have clf.fit train features, y.

And before we run it, let's talk about

this a little bit more. CLF.fit.

So, we're fitting, we're training it. We

are actually creating our random forest

classifier right here. This is the code

that does everything. And we're going to

take our training set. Remember, we kept

our test off to the side. And we're

going to take our training set with the

features. And then we're going to go

ahead and put that in. And here's our

target, the Y. So, the Y is 0, 1, and

two that we just created. And the

features is the actual data going in

that we put into the training set. And

let's go ahead and run that.

And this is kind of an interesting thing

because it printed out the random force

classifier

and everything around it. And so when

you're running this in your terminal

window or in a script like this, this

automatically treats this like just like

when we were up here and I typed in y

and it printed out y instead of print y.

This does the same thing. It treats this

as a variable and prints it out. But if

you were actually running your code,

that wouldn't be the case. And what is

printed out is it shows us all the

different variables we can change. And

if we go down here, you can actually see

in jobs equals 2. You can see the random

state equals zero. Those are the two

that we sent in there. You would really

have to dig deep to find out all these

different meanings of all these

different settings on here. Some of them

are self-explanatory if you kind of

think about it a little bit. Like max

features is auto. So all the features

that we're putting in there, it's just

going to automatically take all four of

them. Whatever we send it, it'll take.

Some of them might have so many features

because you're processing words. There

might be like 1.4 million features in

there because you're doing legal

documents and that's how many different

words are in there. At that point, you

probably want to limit the maximum

features that you're going to process.

And leaf nodes, that's the end nodes.

Remember, we had the fruit and we're

talking about the leaf nodes. Like I

said, there's a lot in this. We're

looking at a lot of stuff here. So you

might have uh in this case there's

probably only think three leaf nodes,

maybe four. You might have thousands of

leaf nodes at which point you do need to

put a cap on that and say, "Okay, you

can only go so far and then we're going

to use all of our resources on

processing this." And that really is

what most of these are about is limiting

the process and making sure we don't uh

overwhelm a system. And there's some

other settings in here. Again, we're not

going to go over all of them. Warm start

equals false. or start as if you're

programming it one piece at a time

externally since we're not we're not

going to have like we're not going to

continually to train this particular

learning tree and again like I said

there's a lot of things in here that

you'll want to look up more detail from

the sklearn and if you're digging in

deep and running a major project on here

for today though all we need to do is

fit or train our features and our target

Y. So now we have our training model.

What's next? If we're going to create a

model,

we now need to test it. Remember, we set

aside the test features, test group, 25%

of the data. So let's go ahead and take

this code and let's put it into our uh

script and see what that looks like.

Okay, here we go. And we're going to run

this.

And it's going to come out with a bunch

of zeros, ones, and twos, which

represents the three type of flowers,

the satossa, the virginica, and the

versa color. And what we're putting into

our predict is the test features. And I

always kind of like to know what it is I

am looking at. So, real quick, we're

going to do test

features. And remember, features is an

array

of sele

width, pedal length, pedal width. So

when we put it in this way, it actually

loads all these different columns that

we loaded into features. So if we did

just features, let me just do features

in here so you can see what features

looks like. This is just playing with

the with Panda's data frames. You'll see

that it's an index. So when you put an

index in like this

into test features into test, it then

takes those columns and creates a Panda

data frames from those columns. And in

this case, we're going to go ahead and

put those into our predict. So, we're

going to put each one of these lines of

data, the 5.0, 3.4, 1.5, point2, and

we're going to put those in, and we're

going to predict what our new um forest

classifier is going to come up with. And

this is what it predicts. It predicts uh

0000121122.

and and uh again this is the flower type

satossa vica and versa color. So now

that we've taken our test features let's

explore that. Let's see exactly what

that data means to us. So the first

thing we can do with our predicts is we

can actually generate a different

prediction model. When I say different,

we're going to view it differently. It's

not that the data itself is different.

So let's take this next piece of code

and put it into our script.

So we're pasting it in here and you'll

see that we're doing uh predict and

we've added underscore proba for

probability. So there's our clff.predict

probability. So we're we're running it

just like we ran it up here, but this

time with this we're going to get a

slightly different result and we're only

going to look at the first 10. So you'll

see down here instead of looking at all

of them uh which was uh what 27 you'll

see right down here that this generates

a much larger field on the probability

and let's take a look and see what that

looks like and what that means. So when

we do the predict underscore probaba for

probability it generates three numbers.

So we had three leaf nodes at the end

and if you remember from all the theory

we did this is the predictors. The first

one is predicting a one for satossa. It

predicts a zero for virginica. And it

predicts a zero for versol. And so on

and so on and so on. And let's um you

know what? I'm going to change this just

a little bit. Let's look at 10

to 20 just because we can.

And we start to get in a little

different of data. And you'll see right

down here it gets to this one. This line

right here. And this line has zero 0.5

0.5.

And so if we're going to vote and we

have two equal votes, it's going to go

with the first one. So it says uh

Satossa gets zero votes, virginica

gets.5 votes, VersaColor gets.5 votes,

but let's just go with the virginica

since these two are equal and so on and

so on down the list. You can see how

they vary on here. So now we've looked

at both how to do a basic predict of the

features and we've looked at the predict

probability. Let's see what's next on

here. So now we want to go ahead and

start mapping names for the plants. We

want to attach names so that it makes a

little more sense for us. And that's

what we're going to do in these next two

steps. We're going to start by setting

up our predictions and mapping them to

the name. So let's see what that looks

like. And let's go ahead and paste that

code in here and run it. And this goes

along with the next piece of code. So

we'll skip through this quickly and then

come back to it a little bit. So, here's

iris.target

names.

And uh if you remember correctly, this

was the the names that we've been

talking about this whole time, the

Satossa, Vica, VersaColor. And then

we're going to go ahead and do the

prediction again. We've run it. We could

have just set a variable equal to this

instead of rerunning it each time, but

we're going ahead and run it again.

CLF.predict test features. Remember that

returns the zeros, the ones, and the

twos. And then we're going to set that

equal to predictions. So this time we're

actually putting it in a variable. And

when I run this,

it distributes and it comes out as an

array. And the array is satossa,

satossa, satossa, satossa, satossa.

We're only looking at the first five. We

could actually do let's do the first 25

just so we can see a little bit more on

there. And you'll see that it starts

mapping it to all the different flower

types, the versa color and the virginica

in there. And let's see how this goes

with the next one. So, let's take a look

at the top part of our species in here.

And we'll take this code and put it in

our script.

And let's put that down here and paste

it. There we go. And we'll go ahead and

run it. And let's talk about both these

sections of code here and how they go

together. The first one is our

predictions. And I went ahead and did uh

predictions through 25. Let's just do

five.

And so we have stosis, satossis, stosis,

satossis. That's what we're predicting

from our test model. And then we come

down here and we look at test species.

And remember, I could have just done

test.species.head.

And you'll see it says Satossa, Satossa,

Satossa, Satossa. And they match. So the

first one is what our forest is doing

and the second one is what the actual

data is. Now is we need to combine these

so that we can understand what that

means. We need to know how good our

forest is, how good it is at predicting

the features. So that's where we come up

to the next step, which is lots of fun.

We're going to use a single line of code

to combine our predictions and our

actuals so we have a nice chart to look

at. And let's go ahead and put that in

our script in our Jupyter notebook here.

Let's see. Let's go ahead and paste that

in. And then I'm going to because I'm on

the Jupyter notebook, I can do a control

minus so we can see the whole line

there.

There we go. resize it and let's take a

look and see what's going on here. We're

going to create in pandas. Remember PD

stands for pandas and we're doing a

cross tab. This function takes two sets

of data and creates a chart out of them.

So when I run it, you'll get a nice

chart down here. And we have the

predicted species.

So across the top you'll see the satossa

versus color virginica and the actual

species satossa versus color virginica.

And so the way to read this chart and

let's go ahead and take a look on how to

read this chart here. When you read this

chart, you have satossa where they meet,

you have versolar where they meet, and

you have virginica where they meet. And

they're meeting where the actual and the

predicted agree. So this is the number

of accurate predictions. So in this

case, it equals 30. If you add 13 + 5 +

12, you get 30. And then we notice here

where it says virginica, but it was

supposed to be versol. This is

inaccurate. So now we have two two

inaccurate predictions and 30 accurate

predictions. So we'll say that the model

accuracy is 93. That's just 30 divided

by 32. And if we multiply it by 100, we

can say that it is 93% accurate. So we

have a 93% accuracy with our model. I

did want to add one more quick thing in

here on our scripting before we wrap it

up. So let's flip back on over to my

script. in here. We're going to take

this uh line of code from up above. I

don't know if you remember it, but

predicts equals the iris.target_names.

So, we're going to map it to the names

and we're going to run the prediction.

And we read it on test features. But,

you know, we're not just testing it. We

want to actually deploy it. So, at this

point, I would go ahead and change this.

And this is an array of arrays. This is

really important when you're running

these to know that. So, you need the

double brackets. And I could actually

create data. Maybe let's let's just do

two flowers. So maybe I'm processing

more data coming in. And we'll put two

flowers in here. And then uh I actually

want to see what the answer is. So let's

go ahead and type in PRS and print that

out. And when I run this, you'll see

that I've now predicted two flowers that

maybe I measured in my front yard as

VersaColor and VersaColor.

Not surprising since I put the same data

in for each one. This would be the

actual uh end product going out to be

used on data that you don't know the

answer for.

So that's going to conclude our

scripting part of this. Introducing

naive base classifier. Have you ever

wondered how your mail provider

implements spam filtering or how online

news channels perform news text

classification or how companies perform

sentimental analysis of their audience

on social media? All of this and more is

done through a machine learning

algorithm called naive bay classifier.

Welcome to Naive Bay tutorial. My name

is Richard Kersner. I'm with the

SimplyLearn team. That's

www.simplearn.com.

Get certified get ahead. What's in it

for you? We'll start with what is naive

bays? A basic overview of how it works.

We'll get into naive bays and machine

learning where it fits in with our other

machine learning tools. Why do we need

naive bays and understanding naive bays

classifier a much more in-depth of how

the math works in the background?

Finally, we'll get into the advantages

of the naive bay classifier in the

machine learning setup. And then we'll

roll up our sleeves and do my favorite

part. We'll actually do some Python

coding and do some text classification

using the naive bays. What is naive

bays? Let's start with a basic

introduction to the bay theorem named

after Thomas Bae from the 1700s who

first coined this in the western

literature. Naive bay classifier works

on the principle of conditional

probability as given by the bay theorem.

Before we move ahead, let us go through

some of the simple concepts in the

probability that we will be using. Let

us consider the following example of

tossing two coins. Here we have two

quarters and if we look at all the

different possibilities of what they can

come up as, we get that they could come

up as head heads. come up as head, tail,

tail, head and tell tail. When doing the

math on probability, we usually denote

probability as a P, a capital P. So the

probability of getting two heads equals

1/4. You can see in our data set, we

have two heads and this occurs once out

of the four possibilities. And then the

probability of at least one tail occurs

three/arters of the time. You'll see on

three of the coin tosses, we have tails

in them. And out of four, that's

three/4s. And then the probability of

the second coin being a head given the

first coin is tail is 1/2. And the

probability of getting two heads given

the first coin is a head is 1/2. We'll

demonstrate that in just a minute and

show you how that math works. Now when

we're doing it with two coins, it's easy

to see. But when you have something more

complex, you can see where these pro

these formulas really come in and work.

So the base theorem gives us the

conditional probability of an event A

given another event B has occurred. In

this case, the first coin toss will be B

and the second coin toss A. This could

be confusing because we've actually

reversed the order of them and go from B

to A instead of A to B. You'll see this

a lot when you work in probabilities.

The reason is we're looking for event A,

we want to know what that is. So, we're

going to label that A since that's our

focus. And then given another event B

has occurred. In the Baze theorem, as

you can see on the left, the probability

of A occurring given B has occurred

equals the probability of B occurring

given A has occurred times the

probability of A over the probability of

B. This simple formula can be moved

around just like any algebra formula.

And we could do the probability of A

after given B times probability of B

equals the probability of B given A

times probability of A. You can easily

move that around and multiply it and

divide it out. Let us apply B theorem to

our example. Here we have our two

quarters and we'll notice that the first

two probabilities of getting two heads

and at least one tail we compute

directly off the data. So you can easily

see that we have one example hh out of

four 1/4 and we have three with tails in

them giving us three quarters or 3/4

75%. The second condition the second uh

set three and four we're going to

explore a little bit more in detail.

Now, we stick to a simple example with

two coins because you can easily

understand the math. The probability of

throwing a tail doesn't matter what

comes before it. And the same with the

heads. So, it's still going to be 50% or

1/2. But when that come when that

probability gets more complicated, let's

say you have a d6 dice or some other

instance, then this formula really comes

in handy. But let's stick to the simple

example for now. In this sample space,

let A be the event that the second coin

is head and b be the event that the

first coin is tails. Again, we reversed

it because we want to know what the

second event's going to be. So, we're

going to be focusing on A. And we write

that out as the probability of A given

B. And we know this from our formula

that that equals the probability of B

given A times the probability of A over

the probability of B. And when we plug

that in, we plug in the probability of

the first coin being tails given the

second coin is heads and the probability

of the second coin being heads given the

first coin being over the probability of

the first coin being tails. When we plug

that data in and we have the probability

of the first coin being tails given the

second coin is heads times the

probability of the second coin being

heads over the probability of the first

coin being tails. You can see it's a

simple formula to calculate. We have 1/2

* 1/2 over 1/2 or 1/2 =.5 or 1/4. So the

B theorem basically calculates the

conditional probability of the

occurrence of an event based on prior

knowledge of conditions that might be

related to the event. We will explore

this in detail when we take up an

example of online shopping further in

this tutorial. Understanding naive bays

and machine learning. Like with any of

our other machine learning tools, it's

important to understand where the naive

bays fits in the hierarchy. So under the

machine learning, we have supervised

learning and there is other things like

unsupervised learning. There's also

reward system. This falls under the

supervised learning. And then under the

supervised learning, there's

classification. There's also regression.

But we're going to be in the

classification side. And then under

classification is your naive bays. Let's

go ahead and glance into where is naive

bays used. Let's look at some of the use

scenarios for it. As a classifier, we

use it in face recognition. Is this

Cindy or is it not Cindy or whoever? Or

it might be used to identify parts of

the face that they then feed into

another part of the face recognition

program. This is the eye. This is the

nose. This is the mouth. Weather

prediction. Is it going to be rainy or

sunny? Medical recognition. News

prediction. It's also used in medical

diagnosis. We might diagnose somebody as

either as high risk or not as high risk

for cancer or heart disease or other

ailments. And news classification you

look at the Google news and it says well

is this political or is this world news

or a lot of that's all done with the

naive bays. Understanding naive bay

classifier. Now we already went through

a basic understanding with the coins and

the two heads and two tails and head

tail tail heads etc. We're going to do

just a quick review on that and remind

you that the naive bay classifier is

based on the bay theorem which gives a

conditional probability of event A given

event B. And that's where the

probability of A given B equals the

probability of B given A times

probability of A over probability of B.

Remember this is an algebraic function

so we can move these different entities

around. We could multiply by the

probability of B. So it goes to the left

hand side and then we could divide by

the probability of A given B and just as

easily come up with a new formula for

the probability of B. To me staring at

these algebraic functions kind of gives

me a slight headache. It's a lot better

to see if we can actually understand how

this data fits together in a table. And

let's go ahead and start applying it to

some actual data so you can see what

that looks like. So, we're going to

start with the shopping demo problem

statement. And remember, we're going to

solve this first in a table form so you

can see what the math looks like. And

then we're going to solve it in Python.

And in here, we want to predict whether

the person will purchase a product. Are

they going to buy or don't buy? Very

important. If you're running a business,

you want to know how to maximize your

profits or at least maximize the

purchase of the people coming into your

store. And we're going to look at a

specific combination of different

variables. In this case, we're going to

look at the day, the discount, and the

free delivery. And you can see here

under the day we want to know whether

it's uh on the weekday, you know,

somebody's working, they come in after

work or maybe they don't work. Weekend,

you can see the bright colors coming

down there celebrating not being in work

or holiday. And did we offer a discount

that day? Yes or no. Did we offer free

delivery that day? Yes or no. And from

this, we want to know whether the

person's going to buy based on these

traits so we can maximize them and find

out the best system for getting somebody

to come in and purchase our goods and

products from our store. Now, having a

nice visual is great, but we do need to

dig into the data. So, let's go ahead

and take a look at the data set. We have

a small sample data set of 30 rows.

We're showing you the first 15 of those

rows for this demo. Now, the actual data

file you can request. Just type in below

under the comments on the YouTube video

and we'll send you some more information

and send you that file. As you can see

here, the file is very simple columns

and rows. We have the day, the discount,

the free delivery, and did the person

purchase or not. And then we have under

the day whether it was a weekday, a

holiday, was it the weekend? This is a

pretty simple set of data. And long

before computers, people used to look at

this data and calculate this all by

hand. So let's go ahead and walk through

this and see what that looks like when

we put that into tables. Also note in

today's world, we're not usually looking

at three different variables and 30

rows. Nowadays, because we're able to

collect data so much, we're usually

looking at 27, 30 variables across

hundreds of rows. The first thing we

want to do is we're going to take this

data and uh based on the data set

containing our three inputs day,

discount, and free delivery, we're going

to go ahead and populate that to

frequency tables for each attribute. So,

we want to know if they had a discount,

how many people buy and did not buy. Uh

did they have a discount? Yes or no. Do

we have a free delivery? Yes or no. On

those days, how many people made a

purchase and how many people didn't? And

the same with the three days of the

week. Was it a weekday, a weekend, a

holiday? And did they buy? Yes or no? As

we dig in deeper to this table for our

bay theorem, let the event buy be a. Now

remember when we looked at the coins, I

said we really want to know what the

outcome is. Did the person buy or not?

And that's usually event A is what

you're looking for. And the independent

variables, discount, free delivery, and

day be B. So we'll call that probability

of B. Now let us calculate the

likelihood table for one of the

variables. Let's start with day, which

includes weekday, weekend, and holiday.

And let us start by summing all of our

rows. So, we have the uh weekday row,

and out of the weekdays, there's 9 plus

2, so there's 11 weekdays. There's eight

weekend days and 11 holidays. Wow,

that's a lot of holidays. And then we

want to sum up the total number of days.

So, we're looking at a total of 30 days.

Let's start pulling some information

from our chart and see where that takes

us. And when we fill in the chart on the

right, you can see that nine out of 24

purchases are made on the weekday, 7 out

of 24 purchases on the weekend, and

eight out of 24 purchases on a holiday.

And out of all the people who come in,

24 out of 30 purchase. You can also see

how many people do not purchase. On the

weekday, it's two out of six didn't

purchase and so on and so on. We can

also look at the totals and you'll see

on the right, we put together some of

the formulas. The probability of making

a purchase on the weekend comes out 11

out of 30. So out of the 30 people who

came into the store throughout the

weekend, weekday and holiday, 11 of

those purchases were made on the

weekday. And then you can also see the

probability of them not making a

purchase. And this is done for doesn't

matter which day of the week. So we call

that probability of no buy would be 6

over 30 or 0.2. So there's a 20% chance

that they're not going to make a

purchase no matter what day of the week

it is. And finally, we look at the

probability of B if A. In this case,

we're going to look at the probability

of the weekday and not buying. Two of

the no buys were done out of the weekend

out of the six people who did not make

purchases. So when we look at that,

probability of the week day without a

purchase is going to be.33 or 33%. Let's

take a look at this at different

probabilities. And uh based on this

likelihood table, let's go ahead and

calculate conditional probabilities as

below. The first three we just did. The

probability of making a purchase on the

weekday is 11 out of 30 or roughly 36 or

37%

367. The probability of not making a

purchase at all doesn't matter what day

of the week is roughly.2 or 20%. And the

probability of a weekday no purchase is

roughly two out of six. So two out of

six of our no purchases were made on the

weekday. And then finally we take our P

of A. If you looked we've kept the

symbols up there. So we got P of

probability of B, probability of A,

probability of B if A. We should

remember that the probability of A if B

is equal to the first one times the

probability of no per buys over the

probability of the weekday. So we could

calculate it both off the uh table we

created. We can also calculate this by

the formula and we get the.367

which equals or.33

*2 over.367 which equals.179

or roughly uh 17 to 18%. And that'd be

the probability of no purchase done on

the weekday. And this is important

because we can look at this and say as

the probability of buying on the weekday

is more than the probability of not

buying on the weekday, we can conclude

that customers will most likely buy the

product on a weekday. Now, we've kept

our chart simple and we're only looking

at one aspect. So, you should be able to

look at the table and come up with the

same information or the same conclusion.

That should be kind of intuitive at this

point. Next, we can take the same setup.

We have the frequency tables of all

three independent variables. Now we can

construct the likelihood tables for all

three of the variables we're working

with. We can take our day like we did

before. We have weekday, weekend, and

holiday. And we filled in this table.

And then we can come in and also do that

for the discount. Yes or no. Did they

buy? Yes or no. And we fill in that full

table. So now we have our probabilities

for a discount and whether the discount

leads to a purchase or not. And the

probability for free delivery. Does that

lead to a purchase or not? And this is

where it starts getting really exciting.

Let us use these three likelihood tables

to calculate whether a customer will

purchase a product on a specific

combination of day, discount, and free

delivery or not purchase. Here, let us

take a combination of these factors. Day

equals holiday, discount equals yes,

free delivery equals yes. Let's dig

deeper into the math and actually see

what this looks like. And we're going to

start with looking for the probability

of them not purchasing on the following

combinations of days. We are actually

looking for the probability of A equal

no buy. No purchase. And our probability

of B we're going to set equal to is it a

holiday? Did they get a discount? Yes.

And was it a free delivery? Yes. Before

we go further, let's look at the

original equation. the probability of a

if b equals the probability of b given

the condition a and the probability

times probability of a over the

probability of b occurring. Now this is

basic algebra so we can multiply this

information together. So when you see

the probability of a given b in this

case the condition is b c and d or the

three different variables we're looking

at. And when you see the probability of

B, that would be the conditions. We're

actually going to multiply those three

separate conditions out. Probability of

you'll see that in just a second in the

formula times the full probability of A

over the full probability of B. So here

we are back to this and we're going to

have let A equal no purchase. And we're

looking for the probability of B on the

condition A where A sets for three

different things. Remember that equals

the probability of A given the condition

B. And in this case, we just multiply

those three different variables

together. So we have the probability of

the discount times the probability of

free delivery times the probability is

the day equal a holiday. Those are our

three variables of the probability of A

if B. And then that is going to be

multiplied by the probability of them

not making a purchase. And then we want

to divide that by the total

probabilities and they're multiplied

together. So we have the probability of

a discount, the probability of a free

delivery, and the probability of it

being on a holiday. When we plug those

numbers in, we see that one out of six

were no purchase on a discounted day,

two out of six were a no purchase on a

free delivery day, and three out of six

were a no purchase on a holiday. Those

are our three probabilities of A of B

multiplied out. And then that has to be

multiplied by the probability of a no

purchase. And remember the prob

probability of a noby is across all the

data. So that's where we get the 6 out

of 30. We divide that out by the

probability of each category over the

total number. So we get the 20 out of 30

had a discount, 23 out of 30 had a yes

for free delivery, and 11 out of 30 were

on a holiday. We plug all those numbers

in, we get.178.

So in our probability math, we have

a.178

if it's a no-by for a holiday, a

discount, and a free delivery. Let's

turn that around and see what that looks

like if we have a purchase. I promise

this is the last page of math before we

dig into the Python script. So here

we're calculating the probability of the

purchase using the same math we did to

find out if they didn't buy. Now we want

to know if they did buy. And again,

we're going to go by the day equals a

holiday, discount equals yes, free

delivery equals yes, and let a equal

buy. Now, right about now, you might be

asking, why are we doing both

calculations? Why why would we want to

know the no buys and buys for the same

data going in? Well, we're going to show

you that in just a moment, but we have

to have both of those pieces of

information so that we can figure it out

as a percentage as opposed to a

probability equation. And we'll get to

that normalization here in just a

moment. Let's go ahead and walk through

this calculation. And as you can see

here, the probability of A on the

condition of B, B being all three

categories, did we have a discount with

a purchase, did we have a free delivery

with a purchase, and did we is a day

equal to holiday. And when we plug this

all into that formula and multiply it

all out, we get our probability of a

discount, probability of a free

delivery, probability of the day being a

holiday times the overall probability of

it being a purchase divided by again

multiplying the three variables out. The

full probability of there being a

discount, the full probability of being

a free delivery, and the full

probability of there being a day equal

holiday. And that's where we get this 19

over 24 * 21 over 24 * 8 over 24 * the p

of a 24 over 30 divided by the

probability of the discount the free

delivery times the day or 20 over 30 23

over 30 * 11 over 30 and that gives us

our 986.

So what are we going to do with these

two pieces of data we just generated?

Well, let's go ahead and go over them.

We have a probability of purchase

equals.986.

We have a probability of no purchase

equals.178.

So finally we have a conditional

probabilities of purchase on this day.

Let us take that we're going to

normalize it and we're going to take

these probabilities and turn them into

percentages. This is simply done by

taking the sum of probabilities which

equals 98686 plus.178

and that equals the 1.164.

If we divide each probability by the

sum, we get the percentage. And so the

likelihood of a purchase is 84.71%.

And the likelihood of no purchase is

15.29%

given these three different variables.

So it's if it's on a holiday, if it's a

with a discount and has free delivery,

then there's an 84.71%

chance that the customer is going to

come in and make a purchase. Hooray,

they purchased our stuff. We're making

money. If you were owning a shop, that's

like is the bottom line is you want to

make some money so you can keep your

shop open and have a living. Now, I

promised you that we were going to be

finishing up the math here with a few

pages. So, we're going to move on and

we're going to do two steps. The first

step is I want you to understand why you

want to why you want to use the naive

bays. What are the advantages of naive

bays? And then once we understand those

advantages, we just look at that

briefly. Then we're going to dive in and

do some Python coding. Advantages of

naive bay classifier. So let's take a

look at the six advantages of the naive

bay classifier. And we're going to walk

around this lovely wheel. Looks like an

origami folded paper. The first one is

very simple and easy to implement.

Certainly you could walk through the

tables and do this by hand. You got to

be a little careful because the

notations can get confusing. You have

all these different probabilities and I

certainly mess those up as I put them

on, you know, is it on the top or the

bottom? We got to really pay close

attention to that. When you put it into

Python, it's really nice because you

don't have to worry about any of that.

You let the Python handle that, the

Python module. But understanding it, you

can put it on a table and you can easily

see how it works. And it's a simple

algebraic function. It needs less

training data. So if you have smaller

amounts of data, this is great powerful

tool for that. Handles both continuous

and discrete data. It's highly scalable

with number of predictors and data

points. So, as you can see, you can just

keep multiplying different probabilities

in there and you can cover not just

three different variables or sets. You

can now expand this to even more

categories. Number five, it's fast. It

can be used in real time predictions.

This is so important. This is why it's

used in a lot of our predictions on

online shopping carts, uh, referrals,

spam filters, is because there's no time

delay as it has to go through and figure

out a neural network or one of the other

mini setups where you're doing

classification. And certainly there's a

lot of other tools out there in the

machine learning that can handle these,

but most of them are not as fast as the

naive bays. And then finally, it's not

sensitive to irrelevant features. So it

picks up on your different

probabilities. And if you're short on

data on one probability, you can kind of

it automatically adjusts for that. Those

formulas are very automatic. And so you

can still get a very solid

predictability even if you're missing

data or you have overlapping data for

two completely different areas. We see

that a lot in doing census and studying

of people and habits where they might

have one study that covers one aspect

and another one that overlaps and

because the two overlap they can then

predict the unknowns for the group that

they haven't done the second study on or

vice versa. So it's very powerful in

that it is not sensitive to the

irrelevant features and in fact you can

use it to help predict features that

aren't even in there. So now we're down

to my favorite part. We're going to roll

up our sleeves and do some actual

programming. We're going to do the use

case text classification. Now, I would

challenge you to go back and send us a

note on the notes below underneath the

video and request the data for the

shopping cart. So, you can plug that

into Python code and do that on your own

time. So, you can walk through it since

we walk through all the information on

it. But, we're going to do a Python code

doing text classification. Very popular

for doing the naive bays. So, we're

going to use our new tool to perform a

text classification of news headlines

and classify news into different topics

for a news website. As you can see here,

we have a nice image of the Google News

and then related on the right subgroups.

I'm not sure where they actually pulled

the actual data we're going to use from.

It's one of the standard sets, but

certainly this can be used on any of our

news headlines in classification. So,

let's see how it can be done using the

naive base classifier. Now, we're at my

favorite part. We're actually going to

write some Python script, roll up our

sleeves, and we're going to start by

doing our imports. These are very basic

imports, including our news group. And

we'll take a quick glance at the target

names. Then we're going to go ahead and

start training our data set and putting

it together. We'll put together a nice

graph because it's always good to have a

graph to show what's going on. And once

we've trained it and we've shown you a

graph of what's going on, then we're

going to explore how to use it and see

what that looks like. Now I'm going to

open up my favorite editor or inline

editor for Python. You don't have to use

this. You can use whatever your editor

that you like, whatever uh interface IDE

you want. This just happens to be the

Anaconda Jupiter notebook. And I'm going

to paste that first piece of code in

here so we can walk through it. Let's

make it a little bigger on the screen so

you have a nice view of what's going on.

Uh and we're using Python 3, in this

case 3.5. So this would work in any of

your 3X if you have it set up correctly.

should also work in a lot of the 2x. You

just have to make sure all of the the

versions of the modules match your

Python version. And in here, you'll

notice the first line is your percentage

mattplot library in line. Now, three of

these lines of code are all about

plotting the graph. This one lets the

notebook know and is inline setup that

we want the graphs to show up on this

page. Without it, in a notebook like

this, which is an explorer interface, it

won't show up. Now, a lot of IDEs don't

require that. A lot of them, like on if

I'm working on one of my other setups,

it just has a popup and the graph pops

up on there. So, you have a that setup

also. But for this, we want the mattplot

library in line. And then we're going to

import numpy as np. That's number

python, which has a lot of different

formulas in it that we use for both of

our sklearn module. And we also use it

for any of the upper math functions in

python. And it's very common to see that

as NP numpy as NP. The next two lines

are all about our graphing. Remember I

said three of these were about graphing.

Well, we need our mattplot

library.pipplot

as plt. And you'll see that plt is a

very common setup as is the sns and just

like the np. And we're going to import

seabor as sns and we're going to do the

sns set. Now seabor sits on top of

pipplot and it just makes a really nice

heat map. It's really good for heat

maps. And if you're not familiar with

heat maps, that just means we give it a

color scale. The term comes from the

brighter red it is, the hotter it is in

some form of data. And you can set it to

whatever you want. And we'll see that

later on. So those you'll see that those

three lines of code here are just

importing the graph function so we can

graph it. And as a data scientist, you

always want to graph your data and have

some kind of visual. It's really hard

just to shove numbers in front of people

and they look at it and it doesn't mean

anything. And then from the sklearn data

sets, we're going to import the fetch 20

news groups. Very common one for

analyzing tokenizing words and setting

them up and exploring how the words work

and how do you categorize different

things when you're dealing with

documents. And then we set our data

equal to fetch 20 news groups. So our

data variable will have the data in it.

And we're going to go ahead and just

print the target names. data.target

names. And let's see what that looks

like. And you'll see here we have alt

atheism comp graphics composs

windows.mmiscellaneous

and it goes all the way down to talk

politics.mmiscellaneous talk

religion.mmiscellaneous. These are the

categories they've already assigned to

this news group and it's called fetch 20

because you'll see there's I believe

there's 20 different topics in here or

20 different categories as we scroll

down. Now, we've gone through the 20

different categories and we're going to

go ahead and start defining all the

categories and set up our data. So,

we're actually getting here going to go

ahead and get it get the data all set up

and take a look at our data. And let's

move this over to our Jupyter notebook.

And let's see what this code does.

First, we're going to set our

categories. Now, if you noticed up here,

I could have just as easily set this

equal to data.target_names target names

because it's the same thing, but we want

to kind of spell it out for you so you

can see the different categories. It

kind of makes it more visual so you can

see what your data is looking like in

the background. Once we've created the

categories, we're going to open up a

train set. So this training set of data

is going to go into fetch 20 news groups

and it's a subset in there called train

and categories equals categories. So

we're pulling out those categories that

match. And then if you have a train set,

you should also have the testing set. We

have test equals fetch 20 news group

subset equals test and categories equals

categories. Let's go down one size so it

all fits on my screen. There we go. And

just so we can really see what's going

on, let's see what happens when we print

out one part of that data. So it creates

train and under train, it creates train

data. And we're just going to look at

data piece number five. And let's go

ahead and run that and see what that

looks like. And you can see when I print

train.data data number five under train.

It prints out one of the articles. This

is article number five. You can go

through and read it on there. And we can

also go in here and change this to test,

which should look identical because it's

splitting the date up into different

groups. Train and test. And we'll see

test number five is a a different

article, but it's another article in

here. And maybe you're curious and you

want to see just how many articles are

in here. We could do length of train.

data. And if we run that, you'll see

that the training data has 11,314

articles. So, we're not going to go

through all those articles. That's a lot

of articles, but um we can look at one

of them just so you can see what kind of

information is coming out of it and what

we're looking at. And we'll just look at

number five for today. And here we have

it. Rewarding the Second Amendment IDs,

VTT, line 58, lines 58 in article, uh

etc. And you can scroll all the way down

and see all the different parts to

there. Now, we've looked at it and

that's pretty complicated when you look

at one of these articles to try to

figure out how do you weight this. If

you look down here, we have different

words and maybe the word from. Well,

from is probably in all the articles.

So, it's not going to have a lot of

meaning as far as trying to figure out

whether this article fits one of the

categories or not. So, trying to figure

out which category it fits in based on

these words is where the challenge comes

in. Now that we've viewed our data,

we're going to dive in and do the actual

predictions. This is the actual naive

bays. And we're going to throw another

model at you or another module at you

here in just a second. We can't go into

too much detail, but it deals

specifically working with words and text

and what they call tokenizing those

words. So, let's take this code and

let's uh skip on over to our Jupyter

notebook and walk through it. And here

we are in our Jupyter notebook. Let's

paste that in there. And I can run this

code right off the bat. It's not

actually going to display anything yet,

but it has a lot going on in here. So

the top we had the print module from the

earlier one. I didn't know why that was

in there. So we're going to start by

importing our necessary packages. And

from the sklearn features

extraction.ext,

we're going to import TF IDF vectorzer.

I told you we're going to throw a module

at you. We can't go too much into the

math behind this or how it works. You

can look it up. The notation for the

math is usually TF.idf.

And that's just a way of weighing the

words. and it weighs the words based on

how many times are used in a document,

how many times or how many documents

they're used in. And it's a well-used

formula. It's been around for a while.

It's a little confusing to put this in

here. Uh, but let's let them know that

it just goes in there and weights the

different words in the document for us.

That way, we don't have to wait. And if

you put a weight on it, if you remember,

I was talking about that up here

earlier. If these are all emails, they

probably all have the word from in them.

From probably has a very low weight. It

has very little value in telling you

what this document's about. Same with

words like in an article in articles in

cost of un maybe cost might or where

words like criminal weapons destruction

these might have a heavier weight

because they describe a little bit more

what the article is doing. Well, how do

you figure out all those weights in the

different articles? That's what this

module does. That's what the TF

vectorizer is going to do for us. And

then we're going to import our

sklearn.na naive bays and that's our

multinnomial NB multinnomial naive bay

pretty easy to understand that where

that comes from and then finally we have

the skyarn pipeline import make pipeline

now the make pipeline is just a cool

piece of code because we're going to

take the information we get from the TF

vectorizer and we're going to pump that

into the multinnomial NB. So, a pipeline

is just a way of organizing how things

flow. It's used commonly. You probably

already guessed what it is. If you've

done any businesses, they talk about the

sales pipeline. If you're on a work crew

or project manager, you have your

pipeline of information that's going

through or your projects and what has to

be done in what order. That's all this

pipeline is. We're going to take the

TFID vectorzer and then we're going to

push that into the multinnomial inb. Now

we've designated that as the variable

model. We have our pipeline model and

we're going to take that model and this

is just so elegant. This is done in just

a couple lines of code. model.fit and

we're going to fit the data. And first

the train data and then the train

target. Now the train data has the

different articles in it. You can see

the one we were just looking at and the

train.target target is what category

they already categorized that that

particular article as. And what's

happening here is the train data is

going into the TF ID vectorizer. So when

you have one of these articles, it goes

in there, it weights all the words in

there. So there's thousands of words

with different weights on them. I

remember once running a model on this

and I literally had 2.4 million tokens

go into this. So when you're dealing

like large document bases, you can have

a huge number of different words. It

then takes those words, gives them a

weight, and then based on that weight,

based on the words and the weights, and

then puts that into the multinnomial NB.

And once we go into our naive bay, we

want to put the train target in there.

So the train data that's been mapped to

the TFID vectorzer is now going through

the multinnomial NB. And then we're

telling it, well, these are the answers.

These are the answers to the different

documents. So this document that has all

these words with these different weights

from the first part is going to be

whatever category it comes out of. Maybe

it's the um talk show or the article on

religion miscellaneous. Once we fit that

model, we can then take labels and we're

going to set that equal to

model.predict. Most of the sklearn use

the term.predict to let us know that

we've now trained the model and now we

want to get some answers. And we're

going to put our test data in there

because our test data is the stuff we

held off to the side. We didn't train it

on there and we don't know what's going

to come up out of it and we just want to

find out how good our labels are. Do

they match what they should be? Now,

I've already run this through. There's

no actual output to it to show. This is

just setting it all up. This is just

training our model, creating the labels

so we can see how good it is, and then

we move on to the next step to find out

what happened. To do this, we're going

to go ahead and create a confusion

matrix and a heat map. So, the confusion

matrix, which is confusing just by its

very name, is basically going to ask how

confused is our answer. Did it get it

correct or did it miss some things in

there or have some missed labels? And

then we're going to put that on a heat

map so we have some nice colors to look

at to see how that plots out. Let's go

ahead and take this code and see how

that uh take a walk through it and see

what that looks like. So, back to our

Jupyter notebook. I'm going to put the

code in there and let's go ahead and run

that code. Take it just a moment. And

remember, we had the inline. That way,

my graph shows up on the inline here.

And let's walk through the code and then

we'll look at this and see what that

means. So, make it a little bit bigger.

There we go. No reason not to use the

whole screen. Too big. So, we have here

from sklearn metrics import confusion

matrix. And that's just going to

generate a set of data that says I the

prediction was such the actual truth was

either agreed with it or was something

different. And it's going to add up

those numbers so we can take a look and

just see how well it worked. And we're

going to set a variable Matt equal to

confusion matrix. We have our test

target, our test data that was not part

of the training. Very important in data

science, we always keep our test data

separate. Otherwise, it's not a valid

model if we can't properly test it with

new data. And this is the labels we

created from that test data. These are

the ones that we predict it's going to

be. So, we go in and we create our SN

heat map. The SNS is our seaborn which

sits on top of the piplot. So, we create

a SNS.heet map. We take our confusion

matrix and it's going to be uh matt.t.

And then we have other variables that go

into the SNS heat map. We're not going

to go into detail what all the variables

mean. The annotation equals true. That's

what tells it to put the numbers here.

So you have the 166, the one, the 00001.

Format D and C bar equals false have to

do with the uh format. If you take those

out, you'll see that some things

disappear. And then the X tick labels

and the Y tick labels. Those are our

target names. And you can see right

here, that's the alt atheism comp

graphics composs windows.mmiscellaneous.

And then finally we have our plt.xl

label. Remember the SNS or the seabor

sits on top of our mattplot library our

plt. And so we want to just tell it x

label equals a true is is true. The

labels are true. And then the y label is

prediction label. So when we say a true,

this is what it actually is. And the

prediction is what we predicted. And

let's look at this graph because that's

probably a little confusing the way I

rattled through it. And what I'm going

to do is I'm going to go ahead and flip

back to the slides because they have a

black background they put in there that

helps it shine a little bit better so

you can see the graph a little bit

easier. So in reading this graph, what

we want to look at is how the color

scheme has come out. And you'll see a

line right down the middle diagonally

from upper left to bottom right. What

that is is if you look at the labels, we

have our predicted label on the left and

our true label on the right. Those are

the numbers where the prediction and the

true come together. And this is what we

want to see is we want to see those lit

up. That's what that heat map does. As

you can see that it did a good job of

finding those data. And you'll notice

that there's a couple of red spots on

there where it missed. You know, it it's

a little confused when we talk about

talk religion miscellaneous versus talk

politics miscellaneous, social religion

Christian versus alt atheism. It

mislabeled some of those. And those are

very similar topics. so you could

understand why it might mislabel them.

But overall, it did a pretty good job.

If we're going to create these models,

we want to go ahead and be able to use

them. So, let's see what that looks

like. To do this, let's go ahead and

create a definition, a function to run.

And we're going to call this function.

Let me just expand that just a notch

here. There we go. I like mine in big

letters. Predict category. So, we want

to predict the category. We're going to

send it as a string. And then we're

sending it train equals train. We have

our training model. And then we had our

pipeline model equals model. This way we

don't have to resend these variables

each time. The definition knows that

because I said train equals train and I

put the equal for model. And then we're

going to set the prediction equal to the

model.predict s. So it's going to send

whatever string we send to it. It's

going to push that string through the

pipeline, the model pipeline. It's going

to go through and uh tokenize it and put

it through the TF IDF, convert that into

numbers and weights for all the

different documents and words. And then

it'll put that through our naive bay.

And from it, we'll go ahead and get our

prediction. We're going to predict what

value it is. And so we're going to

return train.target names predict of

zero. And remember that the train.target

names, that's just categories. I could

have just as easily put uh categories in

there.predict of zero. So we're taking

the prediction which is a number and

we're converting it to an actual

category. We're converting it from um I

don't know what the actual numbers are.

Let's say zero equals alt atheism. So

we're going to convert that zero to the

word or uh one maybe it equals comp

graphics. So we're going to convert

number one into comp graphics. That's

all that is. And then we got to go ahead

and and then we need to go ahead and run

this. So I load that up. And then once I

run that, we can start doing some

predictions. Let me go ahead and type in

predict category. And let's just do

predict category, Jesus Christ. And it

comes back and says it's social,

religion, Christian. That's pretty good.

Now note, I didn't put print on this.

One of the nice things about the Jupiter

notebook editor and a lot of inline

editors is if you just put the name of

the variable out, it's returning the

variable train.target_ames, target

names. It'll automatically print that

for you. In your own IDE, you might have

to put in print. Let's see where else we

can take this. And maybe you're a space

science buff. So, how about sending load

to international

space station.

And if we run that, we get science

space. Or maybe you're a uh automobile

buff. And let's do um Oh, they were

gonna tell me Audi is better than BMW,

but I'm going to do BMW is better than

an Audi. So maybe our car buff. And we

run that. And you'll see it says

recreational. I'm assuming that's what

RECC stands for. Autos. So I did a

pretty good job labeling that one. How

about uh if we have something like a

caption running through there, President

of India. And if we run that, it comes

up and says talk politics miscellaneous.

So when we take our definition or our

function and we run all these things

through, kudos, we made it. We were able

to correctly classify texts into

different groups based on which category

they belong to using the naive base

classifier. Now we did throw in the

pipeline, the TF IDF vectorzer, we threw

in the graphs. Those are all things that

you don't necessarily have to know to

understand the naive base setup or

classifier, but they're important to

know. One of the main uses for the naive

bays is with the TF IDF tokenizer

vectorzer where it tokenizes a word and

has labels and we use the pipeline

because you need to push all that data

through and it makes it really easy and

fast. You don't have to know those to

understand naive bays but they certainly

help for understanding the industry and

data science. And we can see our

categorizer, our naive base classifier.

We were able to predict the category

religion, space, motorcycles, autos,

politics, and properly classify all

these different things we pushed into

our prediction and our trained model.

Before we dive into the SVM, let's take

a look at applications of the support

vector machine, at least some general

ones that are commonly used with it.

face detection, text and hypertext

categorization, classification of

images, and bioinformatics.

These are only but a few of those that

are used with this SVM. As we go through

this lesson, see if you can figure out

what other ones you could apply it to,

and also what you would want to use some

other tools for. So, in this example,

last week, my son and I visited a fruit

shop. Dad, is that an apple or a

strawberry? So, the question comes up,

what fruit did I just pick up from the

fruit stand? After a couple of seconds,

you can figure out that it was a

strawberry. So, let's take this model a

step further and let's uh why not build

a model which can predict an unknown

data. And in this, we're going to be

looking at some sweet strawberries or

crispy apples. We wanted to be able to

label those two and decide what the

fruit is. And we do that by having data

already put in. So, we already have a

bunch of strawberries. We know our

strawberries and they're already labeled

as such. We already have a bunch of

apples. We know our apples and are

labeled as such. Then once we train our

model, that model then can be given the

new data and the new data is this image.

In this case, you can see a question

mark on it and it comes through and goes

it's a strawberry. In this case, we're

using the support vector machine model.

SVM is a supervised learning method that

looks at data and sorts it into one of

two categories. And in this case, we're

sorting the strawberry into the

strawberry side. At this point, you

should be asking the question, how does

the prediction work? Before we dig into

an example with numbers, let's apply

this to our fruit scenario. We have our

support vector machine. We've taken it

and we've taken labeled sample of data,

strawberries and apples, and we draw on

a line down the middle between the two

groups. This split now allows us to take

new data, in this case an apple and a

strawberry, and place them in the

appropriate group based on which side of

the line they fall in. And that way we

can predict the unknown. As colorful and

tasty as the fruit example is, let's

take a look at another example with some

numbers involved. And we can take a

closer look at how the math works. In

this example, we're going to be

classifying men and women. And we're

going to start with a set of people with

a different height and a different

weight. And to make this work, we'll

have to have a sample data set of female

where we have their height and weight

174, 65, 174, 88, and so on. And we'll

need a sample data set of the male. They

have a height 179, 90, 180 to 80 and so

on. Let's go ahead and put this on a

graph so we have a nice visual. So you

can see here we have two groups based on

the height versus the weight. And on the

left side we're going to have the women,

on the right side we're going to have

the men. Now if we're going to create a

classifier, let's add a new data point

and figure out if it's male or female.

So before we can do that, we need to

split our data first. We can split our

data by choosing any of these lines. In

this case, we draw in two lines through

the data in the middle that separates

the men from the women. But to predict

the gender of a new data point, we

should split the data in the best

possible way. And we say the best

possible way because this line has a

maximum space that separates the two

classes. Here you can see there's a

clear split between the two different

classes. And in this one, there's not so

much a clear split. This doesn't have

the maximum space that separates the

two. That is why this line best splits

the data. We don't want to just do this

by eyeballing it. And before we go

further, we need to add some technical

terms to this. We can also say that the

distance between the points in the line

should be as far as possible. In

technical terms, we can say the distance

between the support vector and the hyper

plane should be as far as possible. And

this is where the support vectors are

the extreme points in the data set. And

if you look at this data set, they have

circled two points which seem to be

right on the outskirts of the women and

one on the outskirts of the men. And

hyper plane has a maximum distance to

the support vectors of any class. Now

you'll see the line down the middle and

we call this the hyper plane because

when you're dealing with multiple

dimensions, it's really not just a line

but a plane of intersections. And you

can see here where the support vectors

have been drawn in dashed lines. The

math behind this is very simple. We take

D+ the shortest distance to the closest

positive point which would be on the

men's side and D minus is the shortest

distance to the closest negative point

which is on the women's side. The sum of

D plus and D minus is called the

distance margin or the distance between

the two support vectors that are shown

in the dashed lines. And then by finding

the largest distance margin, we can get

the optimal hyper plane. Once we've

created an optimal hyper plane, we can

easily see which side the new data fits

in. And based on the hyper plane, we can

say the new data point belongs to the

male gender. Hopefully that's clear how

that works on a visual level. As a data

scientist, you should also be asking

what happens if the hyper plane is not

optimal. If we select a hyper plane

having low margin, then there is a high

chance of mclassification. This

particular SVM model, the one we

discussed so far, is also called

referred to as the LSVM.

So far so clear, but a question should

be coming up. We have our sample data

set. But instead of looking like this,

what if it looked like this where we

have two sets of data, but one of them

occurs in the middle of another set. You

can see here where we have the blue and

the yellow and then blue again on the

other side of our data line. In this

data set, we can't use a hyper plane. So

when you see data like this, it's

necessary to move away from a 1D view of

the data to a two-dimensional view of

the data. And for the transformation, we

use what's called a kernel function. The

kernel function will take the 1D input

and transfer it to a two-dimensional

output. As you can see in this picture

here, the 1D when transferred to a

two-dimensional makes it very easy to

draw a line between the two data sets.

What if we make it even more

complicated? How do we perform an SVM

for this type of data set? Here you can

see we have a two-dimensional data set

where the data is in the middle

surrounded by the green data on the

outside. In this case, we're going to

segregate the two classes. We have our

sample data set and if you draw a line

through, it's obviously not an optimal

hyper plane in there. So to do that, we

need to transfer the 2D to a 3D array.

And when you translate it into a

three-dimensional array using the

kernel, you can see where you can place

a hyper plane right through it and

easily split the data. Before we start

looking at a programming example and

dive into the script, let's look at the

advantage of the support vector machine.

We'll start with highdimensional input

space or sometimes referred to as the

curse of dimensionality. We looked at

earlier one dimension, two dimension,

three dimension. When you get to a

thousand dimensions, a lot of problems

start occurring with most algorithms

that have to be adjusted for. The SVM

automatically does that in

highdimensional space. One of the

highdimensional space, one

highdimensional space that we work on is

sparse document vectors. This is where

we tokenize the words in documents so we

can run our machine learning algorithms

over them. I've seen ones get as high as

2.4 million different tokens. That's a

lot of vectors to look at. And finally,

we have regularization parameter. The

realization parameter or lambda is a

parameter that helps figure out whether

we're going to have a bias or

overfitting of the data. Whether it's

going to be overfitted to very specific

instance or it's going to be biased to a

high or low value. With the SVM, it

naturally avoids the overfitting and

bias problems that we see in many other

algorithms. These three advantages of

the support vector machine make it a

very powerful tool to add to your

repertoire of machine learning tools.

Now, we did promise you a use case

study. We're actually going to dive in

to some Python programming. And so we're

going to go into a problem statement and

start off with the zoo. So in the zoo

example, we have um family members going

to the zoo and we have the young child

going, "Dad, is that a group of

crocodiles or alligators?" Well, that's

hard to differentiate. And zoos are a

great place to start looking at science

and understanding how things work,

especially as a young child. And so we

can see the parents sitting here

thinking, well, what is the difference

between a crocodile and an alligator?

Well, one, crocodiles are larger in

size. Alligators are smaller in size.

Snout width. The crocodiles have a

narrow snout and alligators have a wider

snout. And of course, in the modern day

and age, the father's sitting here is

thinking, "How can I turn this into a

lesson for my son?" And he goes, "Let a

support vector machine segregate the two

groups." I don't know if my dad ever

told me that, but that would be funny.

Now, in this example, we're not going to

use actual measurements and data. We're

just using that for imagery. And that's

very common in a lot of machine learning

algorithms and setting them up. But

let's roll up our sleeves and we'll talk

about that more in just a moment as we

break into our Python script. So here we

arrive in our actual coding and I'm

going to move this into a Python editor

in just a moment. But let's talk a

little bit about what we're going to

cover. First, we're going to cover in

the code the setup, how to actually

create our SVM. And you're going to find

that there's only two lines of code that

actually create it. And the rest of it

is done so quick and fast that it's all

here in the first page. and we'll show

you what that looks like as far as our

data because we're going to create some

data. I talked about creating data just

a minute ago. And so we'll get into the

creating data here and you'll see this

nice correction of our two blobs and

we'll go through that in just a second.

And then the second part is we're going

to take this and we're going to bump it

up a notch. We're going to show you what

it looks like behind the scenes. But

let's start with actually creating our

setup. I like to use the Anaconda

Jupyter notebook because it's very easy

to use, but you can use any of your

favorite Python editors or setups and go

in there. But let's go ahead and switch

over there and see what that looks like.

So here we are in the Anaconda Python

notebook or Anaconda Jupyter notebook

with Python. We're using Python 3. I

believe this is 3.5, but it should be

work in any of your 3x versions. And uh

you'd have to look at the sklearn and

make sure if you're using a 2x version,

an earlier version. Let's go and put our

code in there. And one of the things I

like about the Jupyter notebook is I can

go up to view and I'm going to go ahead

and toggle the line numbers on to make

it a little bit easier to talk about.

And we can even increase the size

because this is edited in in this case

I'm using Google Chrome explorer and

that's how it opens up for the editor.

Although anyone any like I said any

editor will work. Now the first step is

going to be our imports and we're going

to import four different parts. The

first two I want you to look at are line

one and line two are numpy as np and

mapplot library.pipplot

as plt. Now these are very standardized

imports when you're doing work. The

first one is the numbers python. We need

that because part of the platform we're

using uses that for the numpy array. And

I'll talk about that in a minute so you

can understand why we want to use a

numpy array versus a standard python

array. And normally it's pretty standard

setup to use NP for numpy. The map plot

library is how we're going to view our

data. So this has uh you do need the NP

for the sklearn module, but the map plot

library is purely for our use for

visualization. And so you really don't

need that for the SVM, but we're going

to put it there so you have a nice

visual aid and we can show you what it

looks like. That's really important at

the end when you finish everything so

you have a nice display for everybody to

look at. And then finally, we're going

to I'm going to jump one ahead to line

number four. That's the sklearn.datas

sets.samples generator import make

blobs. And I told you that we were going

to make up data. And this is a tool

that's in the sklearn to make up data. I

personally don't want to go to the zoo,

get in trouble for jumping over the

fence, and probably get eaten by the

crocodiles or alligators as I work on

measuring their snouts and width and

length. Instead, we're just going to

make up some data. And that's what that

make blobs is. It's a wonderful tool. If

you're ready to test your your uh setup

and you're not sure about what data

you're going to put in there, you can

create this blob and it makes it really

easy to use. And finally, we have our

actual SVM, the sklearn import SVM on

line three. So that covers all our

imports. We're going to create, remember

I used the make blobs to create data.

And we're going to create a capital X

and a lowercase Y equals make blobs in

samples equals 40. So we're going to

make 40 lines of data. It's going to

have two centers with a random state

equals 20. So each each each group's

going to have 20 different pieces of

data in it. And the way that looks is

that we'll have under X um an XY plane.

So I have two numbers under X and Y will

be 01. That's the two different centers.

So we have yes or no in this case

alligator crocodile. That's what that

represents. And then I told you that the

actual sklearn or the SVM is in two

lines of code. And we see it right here

with CLF equals SVM. SVC kernel equals

linear. And I set C equal to one.

Although in this example, since we are

not uh regularizing the data because we

want it to be very clear and easy to

see, I went ahead. You can set it to a

th00and a lot of times when you're not

doing that. But for this thing linear,

because it's a very simple linear

example, we only have the two dimensions

and it'll be a nice linear hyper plane.

It'll be a nice linear line instead of a

full plane. So we're not dealing with a

huge amount of data. And then all we

have to do is do clff.fit

x, y. And that's it. CLF has been

created. And then we're going to go

ahead and display it. And I'm going to

talk about this display here in just a

second. But let me go ahead and run this

code. And this is what we've done is

we've created two blobs. You'll see the

blue on the side and then kind of an

orang-ish uh on the other side. That's

our two sets of data. They represent one

represents crocodiles and one represents

alligators. And then we have our

measurements. In this case, we have like

the width and length of the snout. And I

did say I was going to come up here and

talk just a little bit about our plot.

And you'll see plt. That's what we

imported. We're going to do a scatter

plot. That means we're just putting dots

on there. And then look at this

notation. I have the capital X and then

in brackets I have a colon, 0ero. That's

from numpy. If you did that in a regular

array, you'll get an error in a Python

array. You have to have that in a numpy

array. It turns out that our make blobs

returns a numpy array. And this notation

is great because what it means is the

first part is the colon means we're

going to do all the rows. That's all the

data in our blob we created under

capital X. And then the second part has

a comma 0ero. We're only going to take

the first value. And then if you notice,

we do the same thing, but we're going to

take the second value. Remember, we

always start with zero and then one. So

we have column zero and column one. And

you can look at this as our XY plots.

The first one is the xplot and the

second one is the y plot. So the first

one is on the bottom 0 2 4 6 8 and 10.

And then the second one x of the one is

the 4 5 6 7 8 9 10 going up the left

hand side. S= 30 is just the size of the

dots. We can see them instead of real

tiny dots. And then cmap equals

plt.cm.paired.

And you'll also see the c equals y.

That's the color. We're using two colors

01. And that's why we get the nice blue

and the two different colors for the

alligator and the crocodile. Now you can

see here that we did this the actual fit

was done in two lines of code. A lot of

times there'll be a third line where we

regularize the data. We set it between

like minus one and one and we reshape

it. But for this it's not necessary and

it's also kind of nice because you can

actually see what's going on. And then

if we wanted to we wanted to actually

run a prediction. Let's take a look and

see what that looks like. And to predict

some new data and we'll show this again

as we get towards the end of digging in

deep. You can simply assign your new

data. In this case I am giving it a uh

width and length 34 and a width and

length 56. And note that I put the data

as a set of brackets and then I have the

brackets inside. And the reason I do

that is because when we're looking at

data it's designed to process a large

amount of data coming in. We don't want

to just process one line at a time. And

so in this case, I'm processing two

lines. And then I'm just going to print

and you'll see clf.predict new data. So

the CLF and the predict part is going to

give us an answer. And let's see what

that looks like. And you'll see 01. So

predicted the first one, the 34 is going

to be on the one side and the 56 is

going to be on the other side. So one

came out as a alligator and one came out

as a crocodile. Now that's pretty short

explanation for the setup, but really we

want to dug in and see what it's going

on behind the scenes. and let's see what

that looks like. So, the next step is to

dig in deep and find out what's going on

behind the scenes and also put that in a

nice pretty graph. We're going to spend

more work on this than we did actually

generating the original model. And

you'll see here that we go through a few

steps and I'm I'll move this over to our

editor in just a second. We come in, we

create our original data. It's exactly

identical to the first part and I'll

explain why we redid that and show you

how not to redo that. And then we're

going to go in there and add in those

lines. We're going to see what those

lines look like and how to set those up.

And finally, we're going to plot all

that on here and show it. And you'll get

a nice graph with the what we saw

earlier when we were going through the

theory behind this where it shows the

support vectors and the hyper plane. And

those are done where you can see the

support vectors as the dash lines and

the solid line which is the hyper plane.

Let's get that into our Jupyter

notebook. Before I scroll down to a new

line, I want you to notice line 13. It

has plot show. And we're going to talk

about that here in just a second. But

let's scroll down to a new line down

here. And I'm going to paste that code

in. And you'll see that the plot show

has moved down below. Let's scroll up a

little bit. And if you look at the top

here of our new section, 1 2 3 and four

is the same code we had before. And

let's go back up here and take a look at

that. We're going to fit the values on

our SVM. And then we're going to plot

scatter it. And then we're going to do a

plot show. So you should be asking why

are we redoing the same code. Well, when

you do the plot show, that blanks out

what's in the plot. So once I've done

this plot show, I have to reload that

data. Now, we could do this simply by

removing it up here, rerunning it, and

then coming down here, and then we

wouldn't have to rerun these first four

lines of code. Now, in this, it doesn't

matter too much. And you'll see the plot

show is down here and then removed right

there on line five. I'll go ahead and

just delete that out of there because we

don't want to blank out our screen. We

want to move on to the next setup. So,

we can go ahead and just skip the first

four lines because we did that before.

And let's take a look at the ax=

plt.gca.

Now, right now, we're actually spending

a lot of time just graphing. That's all

we're doing here. Okay. So, this is how

we display a nice graph with our results

and our data. AX is very standard not

used variable when you're talking about

PLT and it's just setting it to that

axis the last axis in the PLT. It can

get very confusing if you're working

with many different layers of data on

the same graph and this makes it very

easy to reference the ax. So this

reference is looking at the PLT that we

created and we already mapped out our

two blobs on. And then we want to know

the limits. So we want to know how big

the graph is. And we can find out the x

limit and the y limit simply with the

get x limit and get yimit commands which

is part of our metplot library. And then

we're going to create a grid. And you'll

see down here we have we've set the

variable xx equal to npines space ximit

0 ximit 1a 30. And we've done the same

thing for the yspace. And then we're

going to go in here and we create a mesh

grid. And this is a numpy command. So

we're back to our numbers python. Let's

go through what these numpy commands

mean with the line space in the mesh

grid. We've taken xx small xx= np line

space. And we have our x limit zero and

our x limit one and we're going to

create 30 points on it. And we're going

to do the same thing for the y axis. Now

this has nothing to do with our

evaluation. It's uh all we're doing is

we're creating a grid of data. And so

we're creating a set of points between

zero and the x limit. We're creating 30

points. And the same thing with the y.

And then the mesh grid loops those all

together. So it forms a nice grid. So if

we were going to do this say between the

limit 0 and 10 and do 10 points, we

would have a 0 0 1 1 0 1 02 03 04 to 10

and so on. You can just imagine a point

at each corner one of those boxes. And

the mesh grid combines them all. So we

take the y and the xx we created and

creates the full grid. And we've set

that grid into the y coordinates and the

xx coordinates. Now remember, when we're

working with Numbi in Python, we like to

separate those. We like to have instead

of it being x comma 1, you know, x comma

y and then x2 comma y2 and in the next

set of data, it would be a column of x's

and a column of y's. And that's what we

have here is we have a column of y's. We

put it as a capital y y and a column of

x's, capital xx with all those different

points being listed. And finally, we get

down to the numpy vstack. Just as we

created those in the mesh grid, we're

now going to put them all into one

array, XY array. Now that we've created

the stack of data points, we're going to

do something interesting here. We're

going to create a value Z. And the Z

equals the CLF. That's our uh that's our

support vector machine we created and

we've already trained. And we have a

decision function. And we're going to

put the XY in there. So here we have all

this data. We're going to put that XY in

there, that data, and we're going to

reshape it. And you'll see that we have

the xx.shape in here. This literally

takes the xx, resets it up, connected to

the y, and the zvalue lets us know

whether it is the left hand side. It's

going to generate three different

values. The zvalue does, and it'll tell

us whether that data is a support vector

to the left, the hyper plane in the

middle, or the support vector to the

right. So it generates three different

values for each of those points. And

those points have been reshaped so

they're right on a line on those three

different lines. So we've set all of our

data up. We've labeled it to three

different areas and we've reshaped it.

And we've just taken 30 points in each

direction. If you do the math, you have

30 * 30. So that's 900 points of data.

And we separated it between the three

lines and reshaped it to fit those three

lines. We can then go back to our map

plot library where we've created the AX

and we're going to create a contour. And

you'll see here where we have contour,

capital XX, capital Y, Y. These have

been reshaped to fit those lines. Z is

the labels. So now we have the three

different points with the labels in

there. And we can set the colors equals

K. And I told you we had three different

labels, but we have uh three levels of

data. The alpha is just makes it kind of

see-through. So it's only uh 0.5 of the

value in there. So when we graph it, the

data will show up from behind it,

wherever the lines go. And finally, the

line styles. This is where we set the

two support vectors to be dash dash

lines and then a single one is just a

straight line. That's what all that

setup does. And then finally, we take

our ax.scatter. We're going to go ahead

and plot the support vectors, but we've

programmed it in there so that they look

nice like the dash dash line and the

dash line on that grid. And you can see

here when we do the CLFS support

vectors, we are looking at column zero

and column one. And then again we have

the S equals 100. So we're going to make

them larger. And the line width equals

1, face colors equals none. Let's take a

look and see what that looks like when

we show it. And you can see when we get

down to our end result, it creates a

really nice graph. We have our two

support vectors and dash lines. And they

have the near data. So you can see those

two points or in this case the four

points where those lines nicely cleave

the data. And then you have your hyper

plane down the middle which is as far

from the two different points as

possible creating the maximum distance.

So you can see that we have our nice

output for the size of the body and the

width of the snout and we've easily

separated the two groups of crocodile

and alligator. Congratulations. You've

done it. We've made it. Of course, these

are pretend data for our crocodiles and

alligators. But this hands-on example

will help you to encounter any support

vector machine projects in the future.

And you can see how easy they are to set

up and look at in depth. We're going to

cover the K nearest neighbors a lot

referred to as KNN. And KNN is really a

fundamental place to start in the

machine learning. It's a basis of a lot

of other things and just the logic

behind it is easy to understand and

incorporated in other forms of machine

learning. So today, what's in it for

you? Why do we need KNN? What is KN&N?

How do we choose the factor K? When do

we use KNN? How does KN&N algorithm

work? And then we'll dive in to my

favorite part, the use case. Predict

whether a person will have diabetes or

not. That is a very common and popular

used data set as far as testing out

models and learning how to use the

different models in machine learning. By

now, we all know machine learning models

make predictions by learning from the

past data available. So we have our

input values. Our machine learning model

builds on those inputs of what we

already know and then we use that to

create a predicted output. Is that a

dog? Little kid looking over there and

watching the black cat cross their path.

No, dear. You can differentiate between

a cat and a dog based on their

characteristics.

Cats. Cats have sharp claws, uses to

climb, smaller length of ears, meows and

purr. Doesn't love to play around. dogs.

They have dull claws, bigger length of

ears, barks, loves to run around. You

usually don't see a cat running around

people, although I do have a cat that

does that where dogs do. And we can look

at these. We can say uh we can evaluate

the sharpness of the claws. How sharp

are their claws? And we can evaluate the

length of the ears. And we can usually

sort out cats from dogs based on even

those two characteristics. Now, tell me

if it is a cat or a dog. Not question.

Usually little kids know cats and dogs

by now. unless you live a place where

there's not many cats or dogs. So, if we

look at the sharpness of the claws, the

length of the ears, and we can see that

the cat has smaller ears and sharper

claws than the other animals. Its

features are more like cats. It must be

a cat. Sharp claws, length of ears, and

it goes in the cat group. Because KN&N

is based on feature similarity, we can

do classification using KN&N classifier.

So, we have our input value, the picture

of the black cat. It goes into our

trained model and it predicts that this

is a cat coming out. So what is knn?

What is the kn&n algorithm? K nearest

neighbors is what that stands for. Is

one of the simplest supervised machine

learning algorithms mostly used for

classification. So we want to know is

this a dog or it's not a dog? Is it a

cat or not a cat? It classifies a data

point based on how its neighbors are

classified. KN&N stores all available

cases and classifies new cases based on

a similarity measure. And here we've

gone from cats and dogs right into wine.

Another favorite of mine. KN&N stores

all available cases and classifies new

cases based on a similarity measure. And

here you see we have a measurement of

sulfur dioxide versus the chloride level

and then the different wines they've

tested and where they fall on that graph

based on how much sulfur dioxide and how

much chloride. K and K&N is a perimeter

that refers to the number of nearest

neighbors to include in the majority of

the voting process. And so if we add a

new glass of wine there, red or white,

we want to know what the neighbors are.

In this case, we're going to put K

equals 5. We'll talk about K in just a

minute. A data point is classified by

the majority of votes from its five

nearest neighbors. Here, the unknown

point would be classified as red since

four out of five neighbors are red. So,

how do we choose K? How do we know K

equals 5? I mean that's was the value we

put in there. I said we're going to talk

about it. How do we choose the factor K?

KN&N algorithm is based on feature

similarity. Choosing the right value of

K is a process called parameter tuning

and is important for better accuracy. So

at K equals 3, we can classify we have a

question mark in the middle as either a

as a square or not. Is it a square or is

it in this case a triangle? And so if we

set K equals to three, we're going to

look at the three nearest neighbors.

We're going to say this is a square. And

if we put k equals a 7, we classify as a

triangle depending on what the other

data is around it. And you can see as

the k changes depending on where that

point is, that drastically changes your

answer. And uh we jump here. We go, how

do we choose the factor of k? You'll

find this in all machine learning.

Choosing these factors, that's the face

you get. It's like, oh my gosh, did I

choose the right K? Did I set it right

my values in whatever machine learning

tool you're looking at? so that you

don't have a huge bias in one direction

or the other. And in terms of KNN, the

number of K, if you choose it too low,

the bias is based on it's just too

noisy. It's it's right next to a couple

things and it's going to pick those

things and you might get a skewed

answer. And if your K is too big, then

it's going to take forever to process.

So you're going to run into processing

issues and resource issues. So what we

do the most common use and there's other

options for choosing k is to use the

square root of n. So n is a total number

of values you have you take the square

root of it. In most cases you also if

it's an even number so if you're using

uh like in this case squares and

triangles if it's even you want to make

your k value odd. That helps it select

better. So in other words you're not

going to have a balance between two

different factors that are equal. So

usually take the square root of n and if

it's even you add one to it or subtract

one from it and that's where you get the

k value from that is the most common use

and it's pretty solid. It works very

well. When do we use kn? We can use kn

when data is labeled. So you need a

label on it. We know we have a group of

pictures with dogs cats cats. Data is

noisefree. And so you can see here when

we have a class and we have like

underweight 140 23 Hello kitty normal

that's pretty confusing. We have a a

high variety of data coming in. So it's

very noisy and that would cause an

issue. Data set is small. So we're

usually working with smaller data sets

where you might get into gig of data if

it's really clean. It doesn't have a lot

of noise because KN&N is a lazy learner.

I.e. it doesn't learn a discriminative

function from the training set. So it's

very lazy. So if you have very

complicated data and you have a large

amount of it, you're not going to use

the kn. But it's really great to get a

place to start. Even with large data,

you can sort out a small sample and get

an idea of what that looks like using

the KN&N and also just using for smaller

data sets. KN&N works really good. How

does the KN&N algorithm work? Consider a

data set having two variables, height in

centimeters and weight in kilograms. And

each point is classified as normal or

underweight. So we can see right here we

have two variables, you know, true

false. They're either normal or they're

not. They're underweight. On the basis

of the given data, we have to classify

the below set as normal or underweight

using KN&N. So if we have new data

coming in that says 57 kg and 177 cm, is

that going to be normal or underweight?

To find the nearest neighbors, we'll

calculate the ukitian distance.

According to the uklitian distance

formula, the distance between two points

in the plane with the coordinates xy and

ab is given by distance d equals the

square root of x - a^2 + y - b^2. And

you can remember that from the two edges

of a triangle. We're computing the third

edge since we know the x side and the y

side. Let's calculate it to understand

clearly. So we have our unknown point

and we placed it there in red. And we

have our other points where the data is

scattered around. The distance d1 is the

square of 170 minus 167^ squar + 57

- 51^ 2ar which is about 6.7 and

distance 2 is about 13 and distance 3 is

about 13.4. Similarly, we will calculate

the ukitian distance of unknown data

point from all the points in the data

set. And because we're dealing with

small amount of data, that's not that

hard to do and it's actually pretty

quick for a computer and it's not a

really complicated math. You can just

see how close is the data based on the

uklidian distance. Hence, we have

calculated the uklidian distance of

unknown data point from all the points

as shown where x1 and y1 equal 57 and

170 whose class we have to classify. So

now we're looking at that. We're saying

well here's the ukitian distance. Who's

going to be their closest neighbors? Now

let's calculate the nearest neighbor at

k equals 3. And we can see the three

closest neighbors puts them at normal.

And that's pretty self-evident when you

look at this graph. It's pretty easy to

say okay what you know we're just voting

normal normal normal. Three votes for

normal. This is going to be a normal

weight. So majority of neighbors are

pointing towards normal. Hence as per

KN&N algorithm the class of 571 170

should be normal. So a recap of KN&N

positive integer K is specified along

with a new sample. We select the K

entries in our database which are

closest to the new sample. We find the

most common classification of these

entries. This is the classification we

give to the new sample. So, as you can

see, it's pretty straightforward. We're

just looking for the closest things that

match what we got. So, let's take a look

and see what that looks like in a use

case in Python. So, let's dive into the

predict diabetes use case. So, use case,

predict diabetes. The objective, predict

whether a person will be diagnosed with

diabetes or not. We have a data set of

768 people who were or were not

diagnosed with diabetes. And let's go

ahead and open that file and just take a

look at that data. And this is in a

simple spreadsheet format. The data

itself is commaepparated. Very common

set of data. And it's also a very common

way to get the data. And you can see

here we have columns A through I. That's

what 1 2 3 4 5 6 7 8. um eight columns

with a particular attribute and then the

ninth column which is the outcome is

whether they have diabetes. As a data

scientist, the first thing you should be

looking at is insulin. Well, you know,

if someone has insulin, they have

diabetes because that's why they're

taking it. And that could cause issue in

some of the machine learning packages,

but for very basic setup, this works

fine for doing the KNN. And the next

thing you notice is it it didn't take

very much to open it up. Um I can scroll

down to the bottom of the data. There's

768.

It's pretty much a small data set. You

know, at 769, I can easily fit this into

my RAM on my computer. I can look at it.

I can manipulate it. And it's not going

to really tax just a regular desktop

computer. You don't even need an

enterprise version to run a lot of this.

So, let's start with importing all the

tools we need. And before that, of

course, we need to discuss what IDE I'm

using. Certainly, you can use any uh

particular editor for Python, but I like

to use for doing uh very basic visual

stuff. the Anaconda, which is great for

doing demos with the Jupyter Notebook.

And just a quick view of the Anaconda

Navigator, which is the new release out

there, which is really nice. You can see

under home, I can choose my application.

We're going to be using Python 3.6. I

have a couple different uh versions on

this particular machine. If I go under

environments, I can create a unique

environment for each one, which is nice.

And there's even a little button there

where I can install different packages.

So, if I click on that button and open

the terminal, I can then use a simple

pip install to install different

packages I'm working with. Let's go

ahead and go back under home and we're

going to launch our notebook. And I've

already, you know, kind of like uh the

old cooking shows, I've already prepared

a lot of my stuff. So, we don't have to

wait for it to launch because it takes a

few minutes for it to open up a browser

window. In this case, I'm going to it's

going to open up Chrome because that's

my default that I use. And since the

script is pre-done, you'll see I have a

number of windows open up at the top,

the one we're working in. And uh since

we're working on the KN&N predict

whether a person will have diabetes or

not. Let's go and put that title in

there. And I'm also going to go up here

and click on cell. Actually, we want to

go ahead and first insert a cell below.

And then I'm going to go back up to the

top cell. And I'm going to change the

cell type to markdown. That means this

is not going to run as Python. It's a

markdown language. So if I run this

first one, it comes up in nice big

letters, which is kind of nice. Remind

us what we're working on. And by now you

should be familiar with doing all of our

imports. We're going to import the

pandas as pd import numpy is np. Pandas

is the uh pandas data frame and numpy is

a number array. Very powerful tools to

use in here. So we have our imports. So

we've brought in our pandas or numpy our

two general python tools. And then you

can see over here we have our train test

split. By now you should be familiar

with splitting the data. We want to

split part of it for training our thing

and then training our particular model

and then we want to go ahead and test

the remaining data to see how good it

is. Pre-processing a standard scaler

pre-processor so we don't have a bias of

really large numbers. Remember in the

data we had like number of pregnancies

isn't going to get very large where the

amount of insulin they take and get up

to 256. So 256 versus six that will skew

results. So we want to go ahead and

change that so they're all uniform

between minus1 and one. And then the

actual tool. This is the K neighbors

classifier we're going to use. And

finally, the last three are three tools

to test. All about testing our model.

How good is it? We just put down test on

there. And we have our confusion matrix,

our F1 score, and our accuracy. So we

have our two general Python modules

we're importing. And then we have our

six modules specific from the sklearn

setup. And then we do need to go ahead

and run this. So these are actually

imported. There we go. And then move on

to the next step. And so in this set,

we're going to go ahead and load the

database. We're going to use pandas.

Remember pandas is pd. And we'll take a

look at the data in Python. We looked at

it in a simple spreadsheet, but usually

I like to also pull it up so that we can

see what we're doing. So here's our data

set equals PD read CSV. That's a pandas

command. And the diabetes folder I just

put in the same folder where my IPython

script is. If you put in a different

folder, you'd need the full length on

there. We can also do a quick length of

uh the data set. That is a simple Python

command. Leen for length. We might even

let's go ahead and print that. We'll go

print. And if you do it on its own line,

length data set in the Jupyter notebook,

it'll automatically print it. But when

you're in most of your different setups,

you want to do the print in front of

there. And then we want to take a look

at the actual data set. And since we're

in pandas, we can simply do data set

head. And again, let's go ahead and add

the print in there. If you put a bunch

of these in a row, you know that data

set one head, data set two head, it only

prints out the last one. So, I usually

always like to keep the print statement

in there. But because most projects only

use one data frame, Panda's data frame,

doing it this way doesn't really matter.

The other way works just fine. And you

can see when we hit the run button, we

have the 768 lines, which we knew, and

we have our pregnancies. It's

automatically given a label on the left.

Remember the head only shows the first

five lines. So we have zero through

four. And just a quick look at the data.

You can see it matches what we looked at

before. We have pregnancy, glucose,

blood pressure all the way to age. And

then the outcome on the end. And we're

going to do a couple things in this next

step. We're going to create a list of

columns where we can't have zero.

There's no such thing as zero skin

thickness or zero blood pressure, zero

glucose. Uh any of those, you'd be dead.

So, not a really good factor if they

don't if they have a zero in there

because they didn't have the data. And

we'll take a look at that because we're

going to start replacing that

information with a couple of different

things. And let's see what that looks

like. So, first we create a nice list.

As you can see, we have the values

talked about glucose, blood pressure,

skin thickness. Uh, and this is a nice

way when you're working with columns is

to list the columns you need to do some

kind of transformation on. Uh, very

common thing to do. And then for this

particular setup, we certainly could use

the there's some Panda tools that will

do a lot of this where we can replace

the NA, but we're going to go ahead and

do it as a data set column equals data

set column.replace. This is this is

still pandas. You can do a direct.

There's also one that that you look for

your nan. A lot of different options in

here. But the nan numpan is what that

stands for is non doesn't exist. So the

first thing we're doing here is we're

replacing the zero with a numpy none.

There's no data there. That's what that

says. That's what this is saying right

here. So put the zero in and we're going

to replace zeros with no data. So if

it's a zero, that means the person's

well hopefully not dead. Hopefully they

just didn't get the data. The next thing

we want to do is we're going to create

the mean which is the in integer from

the data set from the column mean where

we skip NAS. We can do that. That is a

pandas command there, the skip na. So

we're going to figure out the mean of

that data set. And then we're going to

take that data set column and we're

going to replace all the npnan

with the means. Why did we do that? And

we could have actually just uh taken

this step and gone right down here and

just replace zero and skip anything

where except you could actually there's

a way to skip zeros and then just

replace all the zeros. But in this case,

we want to go ahead and do it this way.

So you could see that we're switching

this to a non-existent value. Then we're

going to create the mean. Well, this is

the average person. So if we don't know

what it is, if they did not get the data

and the data is missing, one of the

tricks is you replace it with the

average. What is the most common data

for that? This way you can still use the

rest of those values to do your

computation and it kind of just brings

that particular value or those missing

values out of the equation. Let's go

ahead and take this and we'll go ahead

and run it. Doesn't actually do

anything. So we're still preparing our

data. If you want to see what that looks

like, we don't have anything in the

first few lines, so it's not going to

show up. But we certainly could look at

a row. Let's do that. Let's go into our

data set. Let's print a data set. And

let's pick in this case, let's just do

glucose. And if I run this, this is

going to print all the different glucose

levels going down. And we thankfully

don't see anything in here that looks

like missing data, at least on the ones

it shows. You can see it skipped a bunch

in the middle because that's what it

does. If you have too many lines in

Jupyter notebook, it'll skip a few and

and go on to the next in a data set. Let

me go and remove this. And we'll just

zero out that. And of course, before we

do any processing, before proceeding any

further, we need to split the data set

into our train and testing data. That

way, we have something to train it with

and something to test it on. And you're

going to notice we did a little

something here with the uh pandas

database code. There we go. My drawing

tool. We've added in this right here off

the data set. And what this says is that

the first one in pandas, this is from

the PD pandas. It's going to say within

the data set, we want to look at the eye

location and it is all rows. That's what

that says. So we're going to keep all

the rows, but we're only looking at

zero, column 0 to 8. Remember column 9.

Here it is right up here. We printed it

in here is outcome. Well, that's not

part of the training data. That's part

of the answer. Yeah, it's column 9, but

it's listed as eight. Number eight. So 0

to eight is nine columns. So uh eight is

the value. And when you see it in here,

zero, this is actually 0 to 7. It

doesn't include the last one. And then

we go down here to Y, which is our

answer. And we want just the last one,

just column 8. And you can do it this

way with this particular notation. And

then if you remember, we imported the

train test split that's part of the

sklearn right there. And we simply put

in our X and our Y. We're going to do

random state equals zero. You don't have

to necessarily seed it. That's a seed

number. I think the default is one when

you seated it. I'd have to look that up.

And then the test size. Test size is

0.2. That simply means we're going to

take 20% of the data and put it aside so

that we can test it later. That's all

that is. And again, we're going to run

it. Not very exciting. So far, we

haven't had any print out other than to

look at the data. But that is a lot of

this is prepping this data. Once you

prep it, the actual lines of code are

quick and easy. And we're almost there.

But the actual writing of our KN&N, we

need to go ahead and do a scale the

data. If you remember correctly, we're

fitting the data in a standard scaler,

which means instead of the data being

from, you know, five to 303 in one

column and the next column is 1 to six,

we're going to set that all so that all

the data is between minus1 and one.

That's what that standard scaler does.

Keeps it standardized. And we only want

to fit the scaler with the training set,

but we want to make sure the testing set

is the X test going in is also

transformed. So it's processing it the

same. So here we go with our standard

scaler. We're going to call it sc__x for

the scaler. And we're going to import

the standard scaler into this variable.

And then our xrain equals sc_x.fit

transform. So we're creating the scaler

on the x-ra variable. And then our x

test, we're also going to transform it.

So we've trained and transformed the

x-ra. And then the x test isn't part of

that training. It isn't part of that of

training the transformer. it just gets

transformed. That's all it does. And

again, we're going to go and run this.

And if you look at this, we've now gone

through these steps, all three of them.

We've taken care of replacing our zeros

for key columns that shouldn't be zero,

and we've replaced that with the means

of those columns. That way, that they

fit right in with our data models. We've

come down here, and we split the data.

So, now we have our test data and our

training data. And then we've taken and

we've scaled the data. So all of our

data going in. No, no, we don't tra we

don't train the Y part, the Y train and

Y test that never has to be trained.

It's only the data going in. That's what

we want to train in there. Then define

the model using K neighbors classifier

and fit the train data in the model. So

we do all that data prep. And you can

see down here we're only going to have a

couple lines of code where we're

actually building our model and training

it. That's one of the cool things about

Python and how far we've come. It's such

an exciting time to be in machine

learning because there's so many

automated tools. Let's see. Before we do

this, let's do a quick length of and

let's do y. We want let's just do length

of y. And we get 768. And if we import

math, we do math dot square root. Let's

do y train. There we go. It's actually

supposed to be x train. Before we do

this, let's go ahead and do import math

and do math square root length of y

test. And when I run that, we get

12.409.

I want to see show you where this number

comes from. We're about to use 12 is an

even number. So if you know if you're

ever voting on things, remember the

neighbors all vote. Don't want to have

an even number of neighbors voting. So

we want to do something odd. And let's

just take one away. We'll make it 11.

Let me delete this out of here. That's

one of the reasons I love Jupyter

Notebook because you can flip around and

do all kinds of things on the fly. So,

we'll go ahead and put in our

classifier. We're creating our

classifier now and it's going to be the

K neighbors classifier. In neighbors

equal 11. Remember, we did 12 - 1 for

11. So, we have an odd number of

neighbors. P= 2 because we're looking

for is it are they diabetic or not? And

we're using the ukitian metric. There

are other means of measuring the

distance. You could do like square

square means value. There's all kinds of

measure this, but the uklidian is the

most common one and it works quite well.

It's important to evaluate the model.

Let's use the confusion matrix to do

that. And we're going to use the

confusion matrix. Wonderful tool. And

then we'll jump into the F1 score. And

finally, accuracy score, which is

probably the most commonly used quoted

number when you go into a meeting or

something like that. So, let's go ahead

and paste that in there. And we'll set

the CM equal to confusion matrix. Y

test, Y predict. So those are the two

values we're going to put in there. And

let me go ahead and run that and print

it out. And the way you interpret this

is you have the Y predicted, which would

be your title up here. You can do uh

let's just do Predicted

across the top and actual going down.

Actual. It's always hard to to write in

here. Actual. That means that this

column here down the middle, that's the

important column. And it means that our

prediction said 94 and prediction in the

actual agreed on 94 and 32. This number

here, the 13 and the 15, those are what

was wrong. So you could have like three

different if you're looking at this

across three different variables instead

of just two. You'd end up with a third

row down here in the column going down

the middle. So in the first case, we

have the the and I believe the zero is a

94 people who don't have diabetes. The

prediction said that 13 of those people

did have diabetes and were at high risk.

And the 32 that had diabetes had

correct, but our prediction said another

15 out of that 15, it classified as

incorrect. So you can see where that

classification comes in and how that

works on the confusion matrix. Then

we're going to go ahead and print the F1

score. Let me just run that. And you see

we get a 69 in our F1 score. The F1

takes into account both sides of the

balance of false positives where if we

go ahead and just do the accuracy

account and that's what most people

think of is it looks at just how many we

got right out of how many we got wrong.

So a lot of people when you're a data

scientist and you're talking to other

data scientists they're going to ask you

what the F1 score the Fore is. If you're

talking to the general public or the uh

decision makers in the business, they're

going to ask what the accuracy is. And

the accuracy is always better than the

F1 score. But the F1 score is more

telling. It lets us know that there's

more false positives than we would like

on here. But 82% not too bad for a quick

flash look at people's different

statistics and running an sklearn and

running the KNN, the K nearest neighbor

on it. So we have created a model using

KN&N which can predict whether a person

will have diabetes or not or at the very

least whether they should go get a

checkup and have their glucose checked

regularly or not. The print accuracy

score we got the 0818 was pretty close

to what we got and we can pretty much

round that off and just say we have an

accuracy of 80%. Tells us it is a pretty

fair fit in the model.

>> So what is game means clustering? C

means clustering is an unsupervised

learning algorithm. In this case, you

don't have labeled data unlike in

supervised learning. So you have a set

of data and you want to group them and

as the name suggests, you want to put

them into clusters which means objects

that are similar in nature, similar in

characteristics need to be put together.

So that's what K means clustering is all

about. The term K is basically is a

number. So we need to tell the system

how many clusters we need to perform. So

if K is equal to two, there will be two

clusters. If K is equal to three, three

clusters and so on and so forth. That's

what the K stands for. And of course

there is a way of finding out what is

the best or optimum value of K for a

given data. We will look at that. So

that is K means clustering. So let's

take an example. C means clustering is

used in many many scenarios but let's

take an example of cricket the game of

cricket let's say you received data of a

lot of players from maybe all over the

country or all over the world and this

data has information about the runs

scored by the people or by the player

and the wickets taken by the player and

based on this information we need to

cluster this data into two clusters

batsmen and bowlers. So this is an

interesting example. Let's see how we

can perform this. So we have the data

which consists of primarily two

characteristics which is the runs and

the wickets. So the bowlers basically

take wickets and the batsmen score runs.

There will be of course a few bowlers

who can score some runs and similarly

there will be some batsmen who will who

would have taken a few wickets. But with

this information, we want to cluster

this players into batsmen and bowlers.

So how does this work? Let's say this is

how the data is. So there are

information there is information on the

y-axis about the run scored and on the

x-axis about the wickets taken by the

players. So if we do a quick plot, this

is how it would look. And um when we do

the clustering, we need to have the

clusters like shown in the third diagram

out here. We need to have a cluster

which consists of people who have scored

high runs which is basically the

batsmen. And then we need a cluster with

people who have taken a lot of wickets

which is typically the bowlers. There

may be a certain amount of overlap but

we will not talk about it right now. So

with K means clustering we will have

here that means K is equal to two and we

will have two clusters which is batsmen

and bowlers. So how does this work? The

way it works is the first step in K

means clustering is the allocation of

two centroidids randomly. So two points

are assigned as so-called centrids. So

in this case we want two clusters which

means K is equal to two. So two points

have been randomly assigned as centrids.

Keep in mind these points can be

anywhere. There are random points. They

are not initially they are not really

the centroidids. Centr means it's a

central point of a given data set. But

in this case when it starts off it's not

really the centroid. Okay. So these

points though in our presentation here

we have shown them one point closer to

these data points and another closer to

these data points. They can be assigned

randomly anywhere. Okay. So that's the

first step. The next step is to

determine the distance of each of the

data points from each of the randomly

assigned centrids. So for example we

take this point and find the distance

from this centr and the distance from

this cent. This point is taken and the

distance is found from this centroid and

this c and so on and so forth. So for

every point the distance is measured

from both the centroids and then

whichever distance is less that point is

assigned to that centroid. So for

example in this case visually it is very

obvious that all these data points are

assigned to this centroid and all these

data points are assigned to this

centroid and that's what is represented

here in blue color and in this yellow

color. The next step is to actually

determine the central point or the

actual centrid for these two clusters.

So we have this one initial cluster,

this one initial cluster. But as you can

see these points are not really the

centroid. Centroid means it should be

the central position of this data set.

Central position of this data set. So

that is what needs to be determined as

the next step. So the central point of

the actual centrid is determined and the

original randomly allocated centr is

repositioned to the actual centroid of

this new clusters and this process is

actually repeated. Now what might happen

is some of these points may get

reallocated. In our example that is not

happening probably but it may so happen

that the distance is found between each

of these data points once again with

these centroidids. And if there is if it

is required some points may be

reallocated. We will see that in a later

example but for now we will keep it

simple. So this process is continued

till the centrid repositioning stops and

that is our final cluster. So this is

our so after iteration we come to this

position this situation where the

centroid doesn't need any more

repositioning and that means our

algorithm has converged convergence has

occurred and we have the cluster two

clusters we have the clusters with a

centroid. So this process is repeated.

The process of calculating the distance

and repositioning the centrid is

repeated till the repositioning stops

which means that the algorithm has

converged and we have the final cluster

with the data points and the

centroidids. So this is what you're

going to learn from this session. We

will talk about the types of clustering.

What is K means clustering? application

of K means clustering. C means

clustering is done using distance

measure. So we will talk about the

common distance measures and then we

will talk about how K means clustering

works and go into the details of K means

clustering algorithm and then we will

end with a demo and a use case for K

means clustering. So let's begin. First

of all, what are the types of

clustering? There are primarily two

categories of clustering. hierarchical

clustering and then partitional

clustering and each of these categories

are further subdivided into elomerative

and divisive clustering and K means and

fuzzy C means clustering. Let's take a

quick look at what each of these types

of clustering are. In hierarchical

clustering, the clusters have a treelike

structure and hierarchical clustering is

further divided into elomerative and

divisive. Elomemerative clustering is a

bottomup approach. We begin with each

element as a separate cluster and merge

them into successively larger clusters.

So for example, we have A B CDE E F. We

start by combining B and C form one

cluster. D and E form one more. Then we

combine D, E and F one more bigger

cluster and then add BC to that and then

finally A to it. Compared to that

divisive clustering or divisive

clustering is a top- down approach. We

begin with the whole set and proceed to

divide it into successively smaller

clusters. So we have ABCDE E F. We first

take that as a single cluster and then

break it down into A B C D E and F. Then

we have partitional clustering split

into two subtypes. K means clustering

and fuzzy C means. In K means clustering

the objects are divided into the number

of clusters mentioned by the number K.

That's where the K comes from. So if we

say K is equal to two, the objects are

divided into two clusters C1 and C2. And

the way it is done is the features or

characteristics are compared and all

objects having similar characteristics

are clubed together. So that's how K

means clustering is done. We will see it

in more detail as we move forward. And

fuzzy C means is very similar to K means

in the sense that it clubs objects that

have similar characteristics together.

But while in K means clustering two

objects cannot belong to or any object a

single object cannot belong to two

different clusters in C means objects

can belong to more than one cluster. So

that is the primary difference between K

means and fuzzy C means. So what are

some of the applications of K means

clustering? C means clustering is used

in a variety of examples or variety of

business cases in real life starting

from academic performance, diagnostic

systems, search engines and wireless

sensor networks and many more. So let us

take a little deeper look at each of

these examples. Academic performance. So

based on the scores of the students,

students are categorized into A, B, C

and so on. Clustering forms a backbone

of search engines. When a search is

performed, the search results need to be

grouped together. The search engines

very often use clustering to do this.

And similarly, in case of wireless

sensor networks, the clustering

algorithm plays the role of finding the

cluster heads which collects all the

data in its respective cluster. So

clustering especially K means clustering

uses distance measure. So let's take a

look at what is distance measure. So

while these are the different types of

clustering in this video we will focus

on K means clustering. So distance

measure tells how similar some objects

are. So the similarity is measured using

what is known as distance measure and

what are the various types of distance

measures. There is ukidian distance.

There is Manhattan distance. Then we

have squared ukitian distance measure

and cosine distance measure. These are

some of the distance measures supported

by k means clustering. Let's take a look

at each of these. What is ukidian

distance measure? This is nothing but

the distance between two points. So we

have learned in high school how to find

the distance between two points. This is

a little sophisticated formula for that.

But we know a simpler one is square

roo of y2 - y1 square + x2 - x1

square. So this is an extension of that

formula. So that is the ukidian distance

between two points. What is the squared

ukidian distance measure? It's nothing

but the square of the ukidian distance

as the name suggests. So instead of

taking the square root, we leave the

square as it is. And then we have

Manhattan distance measure. In case of

Manhattan distance, it is the sum of the

distances across the x-axis and the

y-axis. And note that we are taking the

absolute value so that the negative

values don't come into play. So that is

the Manhattan distance measure. Then we

have cosine distance measure. In this

case, we take the angle between the two

vectors formed by joining the points

from the origin. So that is the cosine

distance measure. Okay. So that was a

quick overview about the various

distance measures that are supported by

K means. Now let's go and check how

exactly K means clustering works. Okay.

So this is how K means clustering works.

This is like a flowchart of the whole

process. There is a starting point and

then we specify the number of clusters

that we want. Now there are couple of

ways of doing this. We can do by trial

and error. So we specify a certain

number maybe k is equal to 3 or four or

five to start with and then as we

progress we keep changing until we get

the best clusters or there is a

technique called elbow technique whereby

we can determine the value of k. What

should be the best value of k? How many

clusters should be formed? So once we

have the value of K we specify that and

then the system will assign that many

centrids. So it picks randomly that to

start with randomly that many points

that are considered to be the centrids

of these clusters and then it measures

the distance of each of the data points

from these centroidids and assigns those

points to the corresponding centr from

which the distance is minimum. So each

data point will be assigned to the

centrid which is closest to it and

thereby we have k number of initial

clusters. However this is not the final

clusters. The next step it does is for

the new groups for the clusters that

have been formed it calculates the mean

position thereby calculates the new

centroid position. the position of the

centrid moves compared to the randomly

allocated one. So it's an iterative

process. Once again the distance of each

point is measured from this new centroid

point and if required the data points

are reallocated to the new centroidids

and the mean position or the new centrid

is calculated once again. If the centrid

moves then the iteration continues which

means the convergence has not happened.

The clustering has not converged. So as

long as there is a movement of the

centrid this iteration keeps happening.

But once the centrid stops moving which

means that the cluster has converged or

the clustering process has converged

that will be the end result. So now we

have the final position of the centroid

and the data points are allocated

accordingly to the closest centrid. I

know it's a little difficult to

understand from this simple flowchart.

So let's do a little bit of

visualization and see if we can explain

it better. Let's take an example. If we

have a data set for a grocery shop. So

let's say we have a data set for a

grocery shop and now we want to find out

how many clusters this has to be spread

across. So how do we find the optimum

number of clusters? There is a technique

called the elbow method. So when these

clusters are formed, there is a

parameter called within sum of squares.

And the lower this value is, the better

the cluster is. That means all these

points are very close to each other. So

we use this within sum of squares as a

measure to find the optimum number of

clusters that can be formed for a given

data set. So we create clusters or we

let the system create clusters of a

variety of numbers maybe of 10 10

clusters and for each value of K the

within SS is measured and the value of K

which has the least amount of within SS

or WSS that is taken as the optimum

value of K. So this is the diagrammatic

representation. So we have on the y-axis

the within sum of squares or wss and on

the x-axis we have the number of

clusters. So as you can imagine if you

have k is equal to one which means all

the data points are in a single cluster

the within ss value will be very high

because they are probably scattered all

over. The moment you split it into two

there will be a drastic fall in the

within ss value and that's what is

represented here. But then as the value

of K increases the decrease the rate of

decrease will not be so high. It will

continue to decrease but probably the

rate of decrease will not be high. So

that gives us an idea. So from here we

get an idea for example the optimum

value of K should be either two or three

or at the most four but beyond that

increasing the number of clusters is not

dramatically changing the value in WSS

because that pretty much gets

stabilized. Okay. Now that we have got

the value of K and let's assume that

these are our delivery points. The next

step is basically to assign two centrids

randomly. So let's say C1 and C2 are the

centrids assigned randomly. Now the

distance of each location from the

centrid is measured and each point is

assigned to the centrid which is closest

to it. So for example these points are

very obvious that these are closest to

C1 whereas this point is far away from

C2. So these points will be assigned

which are close to C1 will be assigned

to C1 and these points or locations

which are close to C2 will be assigned

to C2. And then so this is the how the

initial grouping is done. This is part

of C1 and this is part of C2. Then the

next step is to calculate the actual

centrid of this data because remember C1

and C2 are not the centrids. They've

been randomly assigned points and only

thing that has been done was the data

points which are closest to them have

been assigned to them. But now in this

step the actual centroid will be

calculated which may be for each of

these data sets somewhere in the middle.

So that's like the main point that will

be calculated and the centr will

actually be positioned or repositioned

there. Same with C2. So the new centroid

for this group is C2. this new position

and C1 is in this new position. Once

again, the distance of each of the data

points is calculated from these

centroids. Now remember, it's not

necessary that the distance still

remains the or each of these data points

still remain in the same group. By

recalculating the distance, it may be

possible that some points get

reallocated like so. You see this? So

this point earlier was closer to C2

because C2 was here. But after

recalculating repositioning it is

observed that this is closer to C1 than

C2. So this is the new grouping. So some

points will be reassigned. And again the

centrid will be calculated and if the

centroid doesn't change so that is a

repetative process, iterative process.

And if the centroid doesn't change once

the centroid stops changing that means

the algorithm has converged and this is

our final cluster with this as the

centroid C1 and C2 as the centroids

these data points as a part of each

cluster. So I hope this helps in

understanding the whole process

iterative process of K means clustering.

So let's take a look at the K means

clustering algorithm. Let's say we have

x1, x2, x3, n number of points as our

inputs and we want to split this into k

clusters or we want to create k

clusters. So the first step is to

randomly pick k points and call them

centroidids. They are not real centrids

because centr is supposed to be a center

point but they are just called centrids.

And we calculate the distance of each

and every input point from each of the

centroidids. So the distance of X1 from

C1 from C2 C3 each of the distances we

calculate and then find out which

distance is the lowest and assign X1 to

that particular random centroid. Repeat

that process for X2. calculate its

distance from each of the centroid C1,

C2, C3 up to CK and find which is the

lowest distance and assign X2 to that

particular centroid. Same with X3 and so

on. So that is the first round of

assignment that is done. Now we have K

groups because there are we have

assigned the value of K. So there are K

centroids and uh so there are K groups.

All these inputs have been split into K

groups. However, remember we picked the

centrids randomly. So they are not real

centrids. So now what we have to do, we

have to calculate the actual centroids

for each of these groups which is like

the mean position which means that the

position of the randomly selected

centrids will now change and they will

be the main positions of these newly

formed K groups. And once that is done,

we once again repeat this process of

calculating the distance. Right? So this

is what we are doing as a part of step

four. We repeat step two and three. So

we again calculate the distance of X1

from the centroid C1, C2, C3 and then

see which is the lowest value and assign

X1 to that. Calculate the distance of X2

from C1, C2, C3 or whatever up to CK and

find whichever is the lowest distance

and assign X2 to that centroid and so

on. In this process there may be some

reassignment. X1 was probably assigned

to cluster C2 and after doing this

calculation maybe now X1 is assigned to

C1. So that kind of reallocation may

happen. So we repeat the steps two and

three till the position of the centrids

don't change or stop changing and that's

when we have convergence. So let's take

a detailed look at at each of these

steps. So we randomly pick K cluster

centers. We call them centroidids

because they are not initially they are

not really the centrids. So we let us

name them C1 C2 up to CK. And then step

two, we assign each data point to the

closest center. So what we do, we

calculate the distance of each X value

from each C value. So the distance

between X1 C1 distance between X1 C2 X1

C3 and then we find which is the lowest

value. Right? That's the minimum value

we find and assign X1 to that particular

centroid. Then we go next to x2. Find

the distance of x2 from c1, x2 from c2,

x2 from c3 and so on up to ck. And then

assign it to the point or to the

centroid which has the lowest value and

so on. So that is step number two. In

step number three, we now find the

actual centr for each group. So what has

happened as a part of step number two?

We now have all the points, all the data

points grouped into K groups because we

we wanted to create K clusters, right?

So we have K groups. Each one may be

having a certain number of input values.

They need not be equally distributed. By

the way, based on the distance, we will

have K groups. But remember the initial

values of the C1 C2 were not really the

centrids of these groups, right? we

assign them randomly. So now in step

three, we actually calculate the centr

of each group which means the original

point which we thought was the centrid

will shift to the new position which is

the actual centrid for each of these

groups. Okay? And we again calculate the

distance. So we go back to step two

which is what we calculate again the

distance of each of these points from

the newly positioned centroidids and if

required we reassign these points to the

new centroidids. So as I said earlier

there may be a reallocation. So we now

have a new set or a new group. We still

have K groups but the number of items

and the actual assignment may be

different from what was in step two

here. Okay, so that might change. Then

we perform step three once again to find

the new centroid of this new group. So

we have again a new set of clusters, new

centroidids and new assignments. We

repeat this step two again. Once again

we find and then it is possible that

after iterating through three or four or

five times the centrid will stop moving

in the sense that when you calculate the

new value of the centrid that will be

same as the original value or there will

be very marginal change. So that is when

we say convergence has occurred and that

is our final cluster. That's the

formation of the final cluster. All

right. So let's see a couple of demos of

uh K means clustering. We will actually

see some live demos in uh Python

notebook using Python notebook. But

before that let's find out what's the

problem that we are trying to solve. The

problem statement is let's say Walmart

wants to open a chain of stores across

the state of Florida and uh it wants to

find the optimal store locations. Now

the issue here is if they open too many

stores close to each other obviously the

they will not make profit but if they if

the stores are too far apart then they

will not have enough sales. So how do

they optimize this? Now for an

organization like Walmart which is an

e-commerce giant they already have the

addresses of their customers in their

database. So they can actually use this

information or this data and use K means

clustering to find the optimal location.

Now before we go into the Python

notebook and show you the live code, I

wanted to take you through very quickly

a summary of the code in the slides and

then we will go into the Python

notebook. So in this block we are

basically importing all the required

libraries like numpy, mattplot lib and

so on and we are loading the data that

is available in the form of let's say

the addresses for simplicity sake we

will just take them as some data points.

Then the next thing we do is quickly do

a scatter plot to see how they are

related to each other with respect to

each other. So in the scatter plot we

see that there are a few distinct groups

already being formed. So you can

actually get an idea about how the

cluster would look and how many clusters

what is the optimal number of clusters

and then starts the actual K means

clustering process. So we will assign

each of these points to the centrids and

then check whether they are the optimal

distance which is the shortest distance

and assign each of the points data

points to the centroidids and then go

through this iterative process till the

whole process converges and finally we

get an output like this. So we have four

distinct clusters and um which is we can

say that this is how the population is

probably distributed across Florida

state and uh these centroidids are like

the location where the store should be

the optimum location where the store

should be. So that's the way we

determine the best locations for the

store and that's how we can help Walmart

find the best locations for their stores

in Florida. So now let's take this into

Python notebook. Let's see how this

looks when we are learning running the

code live. All right. So this is the

code for K means clustering in Jupyter

notebook. We have a few examples here

which we will demonstrate how K means

clustering is used and even there is a

small implementation of K means

clustering as well. Okay. So let's get

started. Okay. So this block is

basically importing the various

libraries that are required like

mattplot lib and numpy and so on and so

forth which would be used as a part of

the code. Then we are going and creating

blobs which are similar to clusters. Now

this is a very neat feature which is

available in scikitlearn. Make blobs is

a nice feature which creates clusters of

data sets. So that's a wonderful

functionality that is readily available

for us to create some test data kind of

thing. Okay. So that's exactly what we

are doing here. We are using make blobs

and we can specify how many clusters we

want. So centers we are mentioning here.

So it will go ahead and so we just

mentioned four. So it will go ahead and

create some test data for us. And this

is how it looks. As you can see visually

also we can figure out that there are

four distinct classes or clusters in

this data set. And that is what make

blobs actually provides. Now from here

onwards we will basically run the

standard K means functionality that is

readily available. So we really don't

have to implement K means itself. The C

means functionality or the the function

is readily available. You just need to

feed the data and we'll create the

clusters. So this is the code for that.

We import k means and then we create an

instance of k means and we specify the

value of k. This n_clusters is the value

of k. Remember K means in K means K is

basically the number of clusters that

you want to create and it is a integer

value. So this is where we are

specifying that. So we have K is equal

to four and so that instance is created.

We take that instance and as with any

other machine learning functionality fit

is what we use the function or the

method rather fit is what we use to

train the model. Here there is no real

training uh kind of thing but that's the

call. Okay. So we are calling fit and

what we are doing here we are just

passing the data. So x has these values

the data that has been created right. So

that is what we are passing here and uh

this will go ahead and create the

clusters and uh then we are using

after doing uh fit we run the predict

which basically assigns for each of

these observations which cluster it

belongs to. All right. So it will name

the clusters. Maybe this is cluster one.

This is two, three and so on. Or will

actually start from zero, cluster 0, 1,

2 and 3 maybe. And then for each of the

observations it will assign based on

which cluster it belongs to it will

assign a value. So that is stored in y_k

means when we call predict that is what

it does. And we can take a quick look at

these uh y_k means or the cluster

numbers that have been assigned for each

observation. So this is the cluster

number assigned for observation one.

Maybe this is for observation two,

observation three and so on. So we have

how many about I think 300 samples

right? So all the 300 samples there are

300 values here. Each of them the

cluster number is given and the cluster

number goes from 0 to three. So there

are four clusters. So the numbers go

from 0 1 2 3. So that's what is seen

here. Okay. Now, so this was a quick

example of generating some dummy data

and then clustering that. Okay. And this

can be applied if you have proper data.

You can just load it up into X for

example here and then run the K. So this

is the central part of the K means

clustering program example. So you

basically create an instance and you

mention how many clusters you want by

specifying this parameter n_clusters and

that is also the value of k and then

pass the data to get the values. Now the

next section of this code is the

implementation of a k means. Now this is

kind of a a rough implementation of the

k means algorithm. So we will just walk

you through I will walk you through the

code uh at each step what it is doing

and then we will see a couple of more

examples of how K means clustering can

be used in maybe some real life examples

real life use cases. All right. So in

this case here what we're doing is

basically implementing K means

clustering and there is a function or a

library calculates for a given two pairs

of points it will calculate the the

distance between them and see which one

is the closest and so on. So this is

like this is pretty much like what K

means does right. So it calculates the

distance of each point or each data set

from predefined centroid and then based

on whichever is the lowest this

particular data point is assigned to

that centroid. So that is basically

available as a standard function and we

will be using that here. So as explained

in the slides the first step that is

done in case of C means clustering is to

randomly assign some centrides. So as a

first step we randomly allocate a couple

of centrids which we call here we're

calling as centers

and then we put this in a loop and we

take it through an iterative process.

For each of the data points, we first

find out using this function pair-wise

distance argument. For each of the

points, we find out which one which

center or which uh randomly selected

centrid is the closest and accordingly

we assign that data or the data point to

that particular centrid or cluster. And

once that is done for all the data

points, we calculate the new centr by

finding out the mean position with the

the center position. Right? So we

calculate the new centroid and then we

check if the new centroid is the

coordinates or the position is the same

as the previous centroid. The positions

we will compare and if it is the same

that means the process has converged. So

remember we do this process till the

centroidids or the centrid doesn't move

anymore right so the centroid gets

relocated each time this reallocation is

done so the moment it doesn't change

anymore the position of the cent doesn't

change anymore we know that convergence

has occurred so till then so you see

here this is like an infinite loop while

true is an infinite loop it only breaks

when the centers are the same the new

center and old center positions are the

name and once that is uh done uh we

return the centers and the labels. Now

of course as explained this is not a

very sophisticated and advanced

implementation very basic implementation

because one of the flaws in this is that

sometimes what happens is the centroid

the position will keep moving but in the

change will be very minor. So in that

case also that is actually convergence

right. So for example the change is 0.1

we can consider that as convergence

otherwise what will happen is this will

either take forever or it will be never

ending. So that's a small flaw here. So

that is something additional checks may

have to be added here. But again as

mentioned this is not the most

sophisticated uh implementation. This is

like a kind of a rough implementation of

the k means clustering. Okay. So if we

execute this code this is what we get as

the output. So this is the definition of

this particular function and then we

call that find_clusters and we pass our

data x and the number of clusters which

is four and if we run that and plot it

this is the output that we get. So this

is of course each cluster is represented

by a different color. So we have a

cluster in green color, yellow color and

so on and so forth. And these big points

here these are the centroidids is the

final position of the centroidids. And

as you can see visually also this

appears like a kind of a center of all

these points here. Right? Similarly this

is like the center of all these points

here and so on. So this is the example

or this is an example of a

implementation of K means clustering and

uh next we will move on to see a couple

of examples of how K means clustering is

used in maybe some real life scenarios

or use cases. In the next example or

demo, we are going to see how we can use

K means clustering to perform color

compression. We will take a couple of

images. So there will be two examples

and uh we will try to use C means

clustering to compress the colors. This

is a common situation in image

processing when you have an image with

millions of uh colors but then you

cannot render it on some devices which

may not have enough memory. Uh so that

is the scenario where where something

like this can be used. So before again

we go into the Python notebook let's

take a look at quickly the the code. As

usual we import the libraries and then

we import the image and uh then we will

flatten it. So the reshaping is

basically we have the image information

is stored in the form of pixels and uh

if the image is like for example 427x

640 and it has three colors. So that's

the overall dimension of the of the

initial image. we just reshape it and um

then feed this to our algorithm and this

will then create clusters of only 16

clusters. So this this colors there are

millions of colors and now we need to

bring it down to 16 colors. So we use k

is equal to 16 and u this is how when we

visualize this is how it looks. There

are these are all about 16 million

possible colors. The input color space

has 16 million possible colors and we

just sub compress it to 16 colors. So

this is how it would look when we

compress it to 16 colors. And this is

how the original image looks. And after

compression to 16 colors, this is how

the new image looks. As you can see,

there is not a lot of information that

has been lost. though the image quality

is definitely reduced a little bit. So

this is an example which we are going to

now see in Python notebook. Let's go

into the Python and once again as always

we will import some libraries and load

this image called flower.jpg.

Okay. So let we'll load that and this is

how it looks. This is the original image

which has I think 16 million colors and

uh this is the shape of this image which

is basically what is the shape is

nothing but the overall size right so

this is 427 pixel by 640 pixel and then

there are three layers which is this

three basically is for RGB which is red

green blue so color image will have that

right so that is the shape of this now

what we need to do is data let's take a

look at how data is looking. So let me

just create a new cell and show you what

is in data. Basically we have captured

this information.

So data is what? Let me just show you

here.

All right. So let's take a look at

China. What are the values in China? And

uh if you see here, this is how the data

is stored. This is nothing but the pixel

values. Okay? So this is like a matrix

and each one has about for for this 427x

640 pixels. All right. So this is how it

looks. Now the issue here is these

values are large. The numbers are large.

So we need to normalize them to between

0 and one. Right? So that's why we will

basically create one more variable which

is data which will contain the values

between 0 and one. And the way to do

that is divide by 255. So we divide

China by 255 and we get the new values

in data. So let's just run this uh piece

of code and this is the shape. So we now

have also yeah what we have done is we

changed using reshape we converted into

the three-dimensional into a

two-dimensional data set. And let us

also take a look at how

let me just insert

probably a cell here and take a look at

how data is looking. All right. So this

is how data is looking and now you see

this is the values are between 0 and

one. Right? So if you earlier noticed in

case of China the values were large

numbers. Now everything is between 0 and

one. This is one of the things we need

to do. All right. So after that the next

thing that we need to do is to visualize

this and uh we can take random set of

maybe 10,000 points and plot it and

check and see how this looks. So let us

just plot this and uh so this is how the

original the color the pixel

distribution is. These are two plots one

is red against green and another is red

against blue and this is the original

distribution of the color. So then what

we will do is we will use K means

clustering to create just 16 clusters

for the various colors and then apply

that to the image. Now what will happen

is since the data is large because there

are millions of colors using regular K

means may be a little time consuming. So

there is another version of K means

which is called mini batch K means. So

we will use that which is which

processes in the overall concept remains

the same but this basically processes it

in smaller batches. That's the only

thing. Okay. So the results will pretty

much be the same. So let's go ahead and

execute this piece of code and also

visualize this so that we can see that

there are the this is how the 16 colors

uh would look. So this is red against

green and this is red against blue.

there is uh quite a bit of similarity

between this original color schema and

the new one. Right? So it doesn't look

very very completely different or

anything like that. Now we apply this

the newly created colors to the image

and uh we can take a look uh how this is

uh looking. Now we can compare both the

images. So this is our original image

and this is our new image. So as you can

see there is not a lot of information

that has been lost. uh it pretty much

looks like the original image. Yes, we

can see that for example here there is a

little bit uh it appears a little

dullish compared to this one right

because uh we kind of took off some of

the finer details of the color but

overall the highle information has been

maintained. At the same time, the main

advantage is that now this can be this

is an image which can be rendered on a

device which may not be that very

sophisticated. Now let's take one more

example with a different image. In the

second example, we will take an image of

the summer palace in China and we repeat

the same process. This is a high

definition color image with millions of

colors and also uh three-dimensional.

Now we will reduce that to 16 colors

using K means clustering. And um we do

the same process like before. We reshape

it and then we cluster the colors to 16

and then we render the image once again.

And we will see that the color the

quality of the image is slightly

deteriorates. As you can see here, this

has much finer details in this which are

probably missing here. But then that's

the compromise because there are some

devices which may not be able to handle

this kind of a high density images. So

let's run this code in Python notebook.

All right. So let's apply the same

technique for another picture which is

uh even more intricate and has probably

much complicated color schema. So this

is the image. Now once again uh we can

take a look at the shape which is 427x

640x3

and this is the new data would look

somewhat like this compared to the

flower image. So we have some new values

here and we will also bring this as you

can see the numbers are much big. So we

will much bigger so we will now have to

scale them down to values between 0 and

one. And that is done by dividing by

255. So let's go ahead and uh do that

and reshape it. Okay. So we get a

two-dimensional matrix and uh we will

then as the next step we will go ahead

and visualize this how it looks the the

16 colors and this is basically how it

would look 16 million colors. And now we

can create the clusters out of this. The

16 K means clusters we will create. So

this is how the distribution of the

pixels would look with 16 colors. And

then we go ahead and uh apply this and

visualize how it is looking for with the

with the new just the 16 color. So once

again, as you can see, this looks much

richer in color, but at the same time,

and this probably doesn't have, as we

can see, it doesn't look as rich as this

one, but nevertheless, the information

is not lost, the shape and all that

stuff. And this can be also rendered on

a slightly a device which is probably

not that sophisticated. Okay, so that's

pretty much it. So we have seen two

examples of how color compression can be

done uh using K means clustering and we

have also seen in the previous examples

of how to implement C means the code to

roughly how to implement C means

clustering and we use some sample data

using blob to just execute the C means

cluster that takes place after data

collection and before statistical

analysis. So before you conduct any

statistical formulas and analysis on the

data and squeeze the data to extract

some valuable insights, the process

which you perform is called as initial

data analysis. Like taking the data from

the source, cleaning the data,

transforming the data into a readable

format and using that readable data to

build some basic charts what exactly is

happening with this particular company,

brand or anything. Let's say I give you

some data from the company. Then you get

some insights of it. How many number of

traffic you received? How many number of

orders you received? What's the sale

that you made in a specific month,

specific quarter or specific year? And

what was the profit? So basic

information which you convert from the

data and create a dashboard. Right? That

is called as initial data analysis. So a

step beyond initial data analysis is

known as the exploratory data analysis.

This is where you perform some

statistics and probability and predict

the future. Right? So let's dive deep

and learn what exactly is exploratory

data analysis. So a simple definition

for exploratory data analysis is as

follows. Exploratory data analysis is a

key step in data analysis process that

helps you identify patterns, outliners

and relationships between variables

before making assumptions. It is not

like you just create a dashboard out of

the initial data analysis and you can

predict the future. No, it's not like

that. You might have to last. You might

have to go through some permutations and

combinations. You might have to check

the seasons. You might have to check the

possibilities, right? During some

particular seasons in the year, let's

say it's Christmas, then you can expect

some good sales. Let's say it's some

festival, it's some special occasion,

you can expect some good sales on the

product, right? and maybe a part of the

year, maybe a part of 10 years, a

decade, right? In a certain period of

time, there might be some reason due to

which the sales of a certain product

were high. So to make sure that your

assumptions to make sure that your

projection of the sales is 100% accurate

or at least 90 to 95% accurate then you

might have to go through the exploratory

data analysis where you make use of

statistics in your data analysis. Now

there are certain steps that you might

have to follow while going through

exploratory data analysis. So following

are the steps. So a first few steps

might be slightly relevant to initial

data analysis like connecting data,

cleaning it, transforming it and loading

it. After that you will import certain

libraries from Python and after that you

read the data what exactly you have in

your data. The number of columns, the

number of rows and if there are any null

values, if there are any uh entries

which are invalid, you might have to

check that, read that and you might have

to check for duplicate entries. It is

possible that one entry might have been

entered by two different people, right?

There might be a duplication. So you

might have to eliminate those

duplications. You might have to check

for some missing values. You might have

to calculate the total number of missing

values from the data set and try to

eliminate them from the calculation

during your exploratory data analysis.

Followed by that you have to do some

model engineering. Followed by that you

might have to do some feature

engineering creating features and then

you will get started with exploratory

data analysis and the at the end you

will generate a projection or a

prediction or give your assumption that

this might happen in the future and you

might have to take action to avoid it or

you might have to take action to

improvise it right so this is how the

steps in exploratory data analysis take

part now let's proceed and start with

our demo on Python's exploratory data

analysis and in this session we will be

using the use case that we discussed

before which happens to be the students

performance data set. So in this

particular data set we will be having

some columns based on physical activity

the distance from home parental

education the subjects the marks they

have scored in the previous exam. The

marks that they have scored in the

previous exam and if they have any

disabilities if they are having any

resources that they require to write the

exams right. So a few parameters the

important parameters that we will be

discussing in this session and followed

by that we will project the future that

how they will you know improve in their

exams and if there is a problem and if

there is a solution to it then we can

implement that solution and help

students to gain better marks in their

exam. So that's the overall use case for

this demonstration. Now let's get

started with our Jupyter notebook. Now

we are on Jupiter notebook. Now let's

get started. So I would like to have a

title for my um notebook. So I'll write

an HTML code for that.

So HTML code uh

and I have uh three apostrophes

and here I would like to write something

in H1. So I want my title to be in H1.

So

style will be

background color

name will be dark blue.

So we let's uh proceed with the simply

dance background which will be t usually

and the color

of text will be orange

and I also want to have the font let's

let me give the font size as 30.

And in the next line, I'd like to have

border

radius. I can give the border radius as

20 pixels

and padding

to be 16 pixels.

Text alignment, I'd like to keep it

center.

There you go. Let's code the Let's close

the H1. And now let's code the uh

background color or border color. So

B style.

So what we can do is we can basically

have this uh code here and what we will

do is we will reuse this particular code

segment because we will be having

multiple HTML codes in this particular

uh workbook which will explain the

results of the analysis that we are

doing. So the way we just run the code

and after that we will be getting some

visualizations and I will be writing

some textual content in XT in an HTML

page so that it will be easier for the

people to understand what's what exactly

is happening here right so the color

will be light blue and now comes the

text we will close this and here we will

write the text as

Python

explorate

data analysis

and we will

break here

the student performance

and here we will close the H1

and lastly we will display this in HTML

code. So basically we missed this

library. So we will be importing from

ipython display import html to display

this particular code. So without this

particular library we cannot display any

HTML codes in our notebook. So we will

quickly run that and we have a title

over here. Now let's proceed with the

next part. Now we will uh use some

libraries like CAD boost and light bgm.

So for that we might have to install the

these libraries. So we will be using pip

install here

pip install cat boost

and control enter to run this particular

code segment or you can also use run and

it's installed and after that we will

also install

light bgm

g sorry not bgm

so it's already installed

Now we will start importing the

libraries that we need. So we will be

needing numpy, panda, seabon, mattplot

lib and we will also import another

special library which is for ignoring

warnings. So we will import warnings and

after that from we will be importing

that warnings from IPython display

import clear output and after that we

will uh tell the jupyter notebook to

import if there are any warnings. So

that code will be warning dot filter

warnings and in the uh brackets we will

write ignore. So uh the basic libraries

which we will be needing are as follows.

import

numpy

as np. Let's quickly copy this and

proceed. Enter. And now we will be

needing pandas as pd.

Enter. And now we will be needing

seabbone

SNS.

And we will also use mattplot lib

py plot

as plt.

And after that we will import warnings

from

I Python

dot display

import

clear

output

warnings

dot filter warnings

ignore.

Now we'll just quickly run this query.

Run. And now we will proceed with

feature engineering.

So you can use a hashtag to ignore that

particular line from execution for

Jupyter notebook.

And here we will be importing import

from skarn

dot impute import

simple computer

from skarn

dot model

selection

import kf fold. There you go. Now let's

quickly run this query.

There you go. Now we will proceed with

modeling and model evaluation. Once we

are done with this then we will directly

input the data into our notebook. So for

modeling the data we will again use a

hash code so that this particular line

will not execute

and import some libraries lit

GBM

as LGB

from light GBM library.

import

LGBM regressor

from CAD boost import

boost regressor. There you go. Let's

quickly run this command.

There you go. Now, lastly, we have one

more task before importing the data that

is model evaluation. Once that is done,

we can proceed with importing the data.

There you go. Let's quickly run it. And

now so far so good. We have uh done the

basic library imports and feature

engineering is done, modeling is done

and model evaluation is also done. Now

we can begin with importing the data. So

we let's also add um the HTML code where

we have done this importing. So here I

will try to add another segment and I

will import the HTML code here which

displays a similar HTML format which

explains what exactly is happening here.

Just a moment I have the code ready.

I'll just paste it here. There you go.

Let's quickly run this so that we have a

HTML code here page here which explains

what exactly is happening here. So we're

importing libraries and also student

data. Now let's import the student data.

For that let's write the query. So we

are importing the data as data frame and

after that we will write pandas read CSV

right and here we will add the location

of the file. Right? So the file is

located in my downloads section. So

let's quickly copy that location and

paste it here. So this is the location

of my file. Let's quickly run it. So

there might be some error. It's okay. We

can resolve it. So in such scenarios

don't have to worry either you can add

an R but even if that doesn't work you

might have to change uh in this case it

worked but in case if it doesn't work

what you can do is uh you can eliminate

uh the r and you can just change this

from uh forward slash to backlash. This

will also help. So this could be worth

it. So this can also work. So these are

the situations where you can use this.

Now let's proceed with some more uh

interesting facts. Let's try to

understand what's going on with our uh

data set. Right? So what you can do is

now the data is stored in df variable as

a data frame. So what you can do is read

this particular data df.shape shape so

that you can understand what's the uh

what's happening with this data. Right?

So it can tell you that there are uh 6

sorry 6,67

rows and 20 columns. Now let's add

another query part and here you can uh

try to see the head right what head

means basically head means uh the column

headers. So what you can do is just

write df dot head

and run. Now you have the column headers

and couple of sample columns

sorry couple of sample rows. Now let's

proceed with uh checking the duplicates

and identifying the total number of

duplicate entries in this particular

data set. So you can write down df dot

duplicated

dot

sum and you will get the number of

duplicates present in this particular

file. So we have zero duplicate entries.

Now let's see if there are any null uh

elements in this particular data frame.

So df dot is null

dot sum.

So these are all functions. Control

enter. There you go. So in the column

teacher quality there are 78 null

entries and in parental educational

level there are 90 null entries and

distance from home there are u 67 null

entries. Now what we can do is uh from

the okay so from df shape we can add a

new cell here and we can write a HTML

code so that the viewer can understand

that we are trying to understand our

data. So let's write a quick code for

that. So let's not waste much time in

just writing the HTML code. So I've got

that HTML code written in a notepad

already. So I'll just paste it over here

and let's quickly run it so that we have

a HTML visibility here. So this was

supposed to be the result. So we will

add it here. So what we will do is

quickly edit this particular content.

We'll just quickly copy this code and

paste it over here.

Change the content from importing

libraries to

reading

student data and we will cut this code

from here and we will paste it in here

so that it get give us some information.

Let's also run this particular code

segment

so that we have trading student data.

There you go. Now the next part of this

session will be about creating a target

variable. So overall target of this

particular data analysis is about exam

score. Right? So we can name our target

variable as exam score. And let's

understand the distribution of this

particular exam score with uh the

variables we have. Now let's write down

plot dot figure.

Figure size should be around 15 comma 9

equals to let's add a bracket here 15

comma 9 or let's keep it as six 9 would

be a little bigger. Now enter now we

will use seabbond library here. SNS dot

count plot

x is equals to data frame cleaned

target. Okay. Uh before cleaned target

we might have to run a few more. Okay.

We did not perform data cleaning so far,

right? So let's proceed with data

cleaning so far. So we found some empty

entries, right? Null entries and we also

found some So here we have 299 rows

which have missing values. So we will

have to remove that. For that uh we

might have to create a new column which

has to be named as not um assigned right

df data frame not assigned which is

equals to df dot drop na. So we will be

dropping the null values here. It's a

function. And here let's print the

values. Print df dot df

na dot shape. So how many number of rows

and columns we have right and after that

let's also try to eliminate the null

values as well dot is null so we don't

have basically we don't have null values

but we have u some illegal entries maybe

some there you go now let's quickly run

this query so there you go so the new

data is about 678

20

so there is Um, okay. We did uh some

mistake here. So, we supposed to add it

as null. N U N L N N N N N N N N N N N N

N N N N N N L N N N N N N N N N N N N N

N N N N N N N N N N N N N N N N N N N N

N N N N N N N U L. Now quickly run this.

So we should not get any errors this

time.

There you go. No errors. So far so good.

Now let's describe the new data set. df

na dot describe. So these are the new

columns and rows that we have. And we

have uh mean, standard variation,

standard deviation, minimum, maximum. So

the scores are split into 25%, 50% and

75%. Which could be based on hours

studied which could be based on

attendance, sleep hours, previous

course, due training sessions,

everything. So uh minimum sleep hours,

maximum sleep hours, 25% of that, 50% of

that, 75% of that. So that is supposed

to be the u describe. Now what we will

do is we have a target variable which is

exam score. Right? Now we will do some

changes to it. We already know in an

exam there will be a threshold value. It

can be 25 marks per exam. It can be 50

marks per exam and it can be 100 marks

per exam. Right? In our situation let's

consider the threshold value is 100

marks. Right? If there is a situation

where marks is entered in a wrong way

right it can if they add if they wanted

to add 11 but by mistake if they added

uh another one right triple one it's not

a right entry right so we will try to

eliminate those kind of uh data

so we will create a new uh data frame

here which is dataf frame cleaned is

equals to dataf frame

not null so We have eliminated the null

values. BF NA. Now we will add our

target variable which is exam score

should be. So let's come out of this and

here we will add it as should be less

than or equal to 100 but not more than

100. Let's use square brackets.

Here we also the format is square

brackets. ing action now enter and

we will describe this particular

data set instead of the FNA we will copy

paste this here now let's run this

okay exam score is not identified let's

quickly check the error and resolve it

yeah so we missed out to add colons here

it's okay not a problem so this was

supposed to be how it is now let's run

this and we will have the answer over

here. So we have the output. Now let's

check the uh head of this particular

clean data set. So we can make use of

the same code here and paste it right

here and instead of describe let's write

head so that we have the header of uh

this particular data set. So we have our

study and everything normal and we will

categorize the data right. So we will

make use of three columns our study

attendance and previous scores and uh

after that we will also make use of

other columns in this particular data

set which happens to be the parental

involvement access to resources sleep

hours ting sessions etc. And now our

target will be the exam score that we

created over here. Right? This exam

score will be our target. And using this

particular exam score target, we will

categorize the data. Okay? We will

categorize the data in terms of uh let's

say uh first class uh second class and

uh pass or something like that. Right?

So if if a student is uh scoring below

64 and uh that is a separate category.

If the score student is scoring equals

to or above 65, that is a different

category. And if the student is scoring

beyond 70, that's a different category.

And uh before we proceed with that,

let's try to add a HTML code before this

particular data set so that we have u an

understanding of what exactly happened

here. So I let's uh I'll just quickly

copy paste this particular code here. So

we will run it and now next we will

describe check the data type column

separate and everything and we will you

know create data type category variables

right now so far so good. Now we will

create the categories.

So num call

equals to

hours studied

comma attendance.

So let's quickly add the data

previous exam scores.

Just a minute. Let's quickly add the

columns. Let me take a while. There you

go. I've added the columns. So we are

considering three different columns. uh

our study attendance and previous scores

for num call and cat call. We are

considering the other columns apart from

the first three and our target value is

exam score. Let's quickly run this.

There you go. And now we will try to

build some visualizations and before

that let's try to uh add some data uh

from in HTML. Let's try to create a

Okay, what we can do is simply copy this

particular HTML file here. We can take

this

and add it here so that we will

understand what exactly is happening

next. And in place of reading student

data, we will write data visualization

for student data.

And we will keep the colors same dark

blue background and u the color for data

visualization will be light blue and

student data will be orange. Let's

quickly run and there we have it. Now

our target variable is exam score. Right

now we will compare this particular

target variable with three other

parameters. So our parameters will be

the following uh as we discussed uh

creating the segregation in data set

right. So first will be distribution of

target variable with other parameters.

So we will create another HTML file for

that right here. Just a minute while I

paste the code for um HTML quickly run

this. There you go. So distribution of

target variable exam score against some

parameters. Now we will write the plot

for that

plot dot figure. So we are going to

consider the size equals to 15 6

big size

is equals to 15 6 the same one that we

considered before. And we will be using

Cbond SNS dot count plot

open bracket. This is a function x is

equals to df

clean. Okay. Uh what we can do is just

quickly take the column name so that we

don't create any mistakes here and we

will paste it over here instead of df.

There you go. Or target.

So our target is exam score,

and we will use the pellet as green

and the plot title will be distribution

of target variable exam score. We can

copy this. Okay, just a minute before

that plt do.

Should be let's use double quotes now.

Copy this and paste it here.

I think semicolons went off. Okay, not a

problem.

It's right here. Let's add a dot as a

full stop. And the next line, if you

need, you can add the full stop. If not

you can ignore plt dot grid true

which equals to major

comma

access is equals to y

comma line style equals to so I want

lines to be hyphen hyphen in this way I

want the lines and comma line width um

let's say 0.5 or 0.7. Let's go with 0.7

line width equals to 0.9 mm. There you

go. Now let's quickly run this query.

There you go. Done. And we have the

first visualization. So here you can see

there are some students which are

scoring 58 59 and you can see maximum

number of students are already scoring

good marks which is under 65 and uh

sorry which is under 70 and above 65 and

there is a good number of students uh

which are also scoring uh above 65 as

well right so sorry 70 70 as well. So we

have now less than or equal to 70 71.

And highest scorer in some situations

there is also 100. If you can see there

is slight growth here. There are a few

students toppers maybe which have

already scored 100 as well. Now we have

the list here. Now we what we need to do

is we need to segregate that is part one

which is less than or equal to 64 which

falls under 65 and another category

which falls in between 65 to 70 and

above 70. So we need to categorize these

three uh datas and segregate them as

bottom 65, top which is above 70 and mid

between 70 to 65. Right now before that

if you want to add an HTML document uh

sorry segment here which explains about

the distribution of target you can also

do that. It's already added here. Now

let's continue. But in case if you want

to uh add uh some data which explains

that we're trying to segregate, you can

also do that. I would like to do that.

Let's quickly uh add that HTML code

here. So what this particular code will

do is it will tell the percentage of

students which are scoring less than 65.

Number of stu uh percentage of students

uh scoring in between 65 and 70 and the

percent of students which are scoring

beyond 70. Right? Let's run this. And

here we have the result. Bottom 21.81%

scores under 64 while 24% scores over 70

and 50% are in between 65 to 69. Right

now let's uh continue with the

segregation part of the data. So for

segregation we will create three

different variables A, B, C. So first A

is equals to length of DF claimed. So

let's copy the column name sorry data

frame name length of DF cleaned inside

the square brackets we'll again add DF

cleaned of target variable which is exam

score

let's also add uh single quotes here

who are scoring in between or um less

than let's start with less than or equal

to 64 we'll not consider is 65 we'll

consider 64 divided by len of df cleaned

target variable exam score single quotes

into 100

which will give us the percentage now

similarly let's just copy and paste this

three more times for B and C. So here

instead of minus we're supposed to add

equals to and another one. So here

equals to and instead of A I will write

B and the last one is C. And instead of

64 we will add 70 here for top and here

we will make some changes. It should be

greater than or equal to

65. So this is the third category A B C

and then we will proceed with printing

the files. So print

the bottom

for the first one which is f of

a

is to do 2f

and we will add the percentage symbol

over here r under

64.

Now we can copy paste the same here and

we can change the variables.

So here we will be adding under over 70.

Lastly in between the ones in between

65 and 70.

There you go. Here we will change the

values from A to B and here A to C.

There you go. And we can quickly run

this query. So it's not 70. It was

supposed to be 69.

There you go. So we forgot to mention

this particular one. Now let's run this.

There you go. So we have 21%

of people who are scoring under 64, 24%

over 70 and 53% are in between average.

So I think the school is focusing on

improving this particular percentage,

reducing this particular percentage and

increasing this particular percentage

and try to eliminate if possible this

particular one which are under 64. So

that is the overall moto I guess. Now so

far so good. Let's now try to remove

infinite values from HTML, right? So

before that, let's add uh this

particular HTML code here. So which

explains what we are trying to do. So we

will first implement the code that

prevents warning about infinite values

during data visualization. And now let's

add the code which will try to eliminate

the uh infinite values. Let's copy this

particular data frame cleaned uh data

frame name here. Now df cleaned

dotreplace

square brackets np dot info

comma

minus np

dot info out of these square brackets

dot np

na

non na values we're trying to eliminate

na values in place of those values you

can write true and after that we will

try to eliminate the null. So if it is

null

dot sum give me the total number of null

values after this. Right? So let's try

to execute that. There you go. So all

the null entries have been removed here.

Now let's see the distribution of

numerical values here. So before that

let's add the HTML code for that. So

let's quickly run this. So distribution

of numerical values. So we will be

considering these three parameters. So

if you go back here you can see our

studied attendance and previous scores.

So we will be considering these three

values or these three columns and check

the distribution of these variables

against the exam score. So uh is it

making any um you know kind of variation

if the u attendance is high? If if the

attendance is high is the exam score

high and uh apart from that we have if

our studies is high is the uh mark score

is high and if the previous scores are

high is there a chance to get better

scores in this particular exam. So what

we are trying to do is we are trying to

see if there is any direct involvement

of number of study hours and number of

days attended and number of uh or the

number of marks they received in the

previous course and we'll try to build a

visualization on that front.

So let's go and build that. So we'll try

the try to write the code here.

Figure axis

equals tot

dot

subplots. So we will be having three

different plots here since we're

considering three different u

categories.

And the fixed size should be equal to

12A 4. There you go. Now

access

is equals to access dot variable

per idx

comma call

in enumerate

and we will try to import the seaborn

library here. We will try to create

histo plots here. Histograms here

plot.

So line style we will be selecting this

one

and the comma

line width will be 0.7.

There you go. Next will be access

dot set

title. So for this we will be uh setting

the title as distribution of columns. So

the columns will be the three uh ones

attendance, hours studied and uh what

was the third one that we considered

previous course. Right? So instead of

mentioning them specifically, what we

can do is we can just write columns

here. C L and close.

There you go. And lastly,

plt.tight Right.

And show the plot. There you go. Let's

quickly run this query. Run. And now we

will be having the visualizations here.

So um you can directly see the

involvement of these three parameters

here. If uh they are trying to help if

the number of hours are increased then

you can see if there is a better

improvement in scores. If the attendance

is increased, if there is a betterment

in scores or if the previous uh scores

are helping then you can find it out how

it is. There you go. Now we can write a

result here in the form of HTML page.

And if we run this, it will give you the

result. The breaks or gaps in the hour

study variable may be due to the

respondents answering appropriately. The

variables attendance and previous scores

which exhibit a uniform distribution

have a normal impact on exam scores

variable which is our target variable.

Now let's proceed with another part of

this session which will be about the

relationship between the numerical

values and the target variables. Now we

will copy paste the same code and make

some minute changes to it. So the only

change that we did to it is we're trying

to uh build a scatter plot. So we will

be getting a scatter plot here. But

before that let's try to add another

column here and try to add an HTML code

which explains why we are doing it.

There you go a scatter plot. So

basically these two are one and the

same. Here we use some column graphs. So

here we did the same using the scatter

plot which will help for a better

understanding. Now we will try to build

some correlations.

So basically a list is called

correlation is created containing the

names of the columns for which you want

to calculate the correlation. So here in

our situation it is the df c r which is

a data frame and it is created by

selecting only these columns for df

cleaned data frame effectively created a

new data frame containing only the

specified columns. Now the second one

which is the co r which is a calculate

correlation. So this method computes the

correlation matrix for selected columns

which is n df c r. The one indicates

perfect positive correlation minus one

indicates the perfect negative

correlation and zero indicates no

correlation. And lastly the setup of

plot. This line sets up the figure size

of the plot. In this particular

situation we are choosing five and four.

Right now let's quickly try to execute

this query and see the answer. And we

will also add the HTML code for this so

that we have a better understanding for

this. So we will be adding that HTML

uh box here which will explain what

exactly happened here. So this is our

plot and this is the correlation. Now

let's try to add that HTML code right

here. The result of this particular data

visualization will be maintained here.

So the hours studying and attendance

shows a positive correlation with the

target variable which is exams hour.

However, previous course appears to have

no or little relationship with the

target variable. Right now let's

continue with our next uh part of this

session. So now we will try to identify

the relationship between studies

hours and attendance and extracurricular

scores. Right? So we have other u

columns to consider which is

extracurricular activities. So there is

a belief that extracurricular activities

will also help students to study better.

So we will find if there is a relation

between the target variable and this

extracurricular activities attendance

and study hours. Now let's quickly add

the code here. Now let's quickly execute

the code. Now we have the visualization

which explains the relationship between

the number of hours studied

extracurricular activities etc. So here

it is and now let's add an HTML code

which explains about this result

in this particular code was supposed to

be added here.

So this uh is the resultant column here.

Now here it explains about the

influences that it performs. So the

extracal activities, parental income and

extra things that influence the scores

and there you go.

Now let us also consider other columns

right the other parameters like

resources are available or not parental

education and other things which also

might have influenced the exam scores of

students. So for that let's add an HTML

code so that we have an HTML page here

which explains what is the next

procedure that we are following. Right

now let's add the query here. So here we

are considering the other parameters

like family income, peer influence,

motivation level, gender, parental

involvement, parental educational level

and extracurricular activities. And we

are considering them against the target

variable which is exam score. And now

let's execute this query. There you go.

Now we have generated a graph which

explains about this particular

parameters against the target variable.

And now let's add the HTML page here

which explains about these results.

Let's quickly run it. And there you go.

So when certain factors affect Q1 and Q2

but not Q2, it can be understood that

individual has overcome challenges

through personal effort. Right? So if

government policies and corporate social

contributors are focused on addressing

these aspect, it seems that we could

create dynamic country with greater

social mobility and open opportunities

for all. Right? So if extracurricular

activities can outweigh the influence of

other variables in academic performance

then we should foster that kind of

environment right. So this is how u you

can get extract some statistical

analysis on this particular data set.

Now let's quickly rename this uh python

eda

students

performance

and you can quickly rename and save it.

Welcome to math refresher probability

and statistics.

In this lesson, we are going to explain

the concepts of statistics and

probability.

Describe conditional probability. Define

the chain rule of probability. Discuss

the measure of variance. Identify the

types of gshian distribution.

Basic of statistics and probability.

Probability and statistics. Data science

relies heavily on estimates and

predictions. A significant portion of

data science is made up of evaluations

and forecast.

Statistical methods are used to make

estimates for further analysis.

Probability theory is helpful for making

predictions. Statistical methods are

highly dependent on probability theory

and all probability and statistics are

dependent on data. Data is information

acquired for reference or research via

observations, facts, and measurements.

Data is a set of facts structured in the

form that computers can interpret such

as numbers, words, estimations, and

views. Importance of data. Data aids in

seeing more about the information by

identifying possible connections between

two features. Data assists in the

detection of distortion by uncovering

hidden patterns based on prior

information patterns. Data may be

utilized to anticipate the future or

predict the current state of affairs.

Also, data aids in determining whether

two pieces of information have any

instance in common or not. Types of

data. Data might be quantitative. That

is data that can be measured or counted

in numbers. Or it may be qualitative

which is data which is generally divided

into groups or in simpler words which

cannot be counted or measured in

numbers. Let's consider an example. A

customer information data of a bank may

contain quantitative and qualitative

data. Consider this snapshot where we

have customer ID, surname, geography,

gender, age, balance, has C or card is

active member. Amongst these variables

we can see surname is mostly qualitative

as it cannot be counted and measured in

numbers. Geography and gender are also

qualitative as they cannot be counted in

numbers and are mostly groups. has C or

card that is has credit card and is

active member although are containing

numerical in form but these are

categorical that means these have been

divided into groups of one and zero that

represent yes and no as an answer hence

these two variables are also qualitative

customer ID is again although a

numerical data however the significance

or intuition behind Customer ID is

categorical.

Hence, it may be kept in the qualitative

data also. However, age and balance

these are numerical information which

have been measured or counted and

numerical operations can be performed on

them. Hence, these are under

quantitative data categories.

Introduction to descriptive statistics.

Descriptive statistics. A descriptive

measurement is summary measure that

quantitatively portrays the most

important features of a set of data

allowing for a better comprehension of

the information. Data can be measured as

different levels. The levels of

measurement describe the nature of

information stored in the data assigned

to the variables. Qualitative data can

be measured as nominal or ordinal.

Quantitative data can be measured in

terms of interval and ratio type.

Nominal data. The data is categorized

using names, labels or qualities. For

example, brand name, zip code, and

gender. Ordinal data can be arranged in

order or ranked, and can be compared.

Examples include grades, star reviews,

position, and race, and date. Interval

data is the data that is ordered and has

meaningful differences between the data

points. Example, temperature in Celsius

and year of birth. Ratio data is similar

to the interval level with the added

property of inherent zero. Mathematical

calculations can be performed on both

interval as well as ratio data. For

example, height, age, and weight.

Population versus sample. Before

analyzing the data, it's important to

figure out if it's from a population or

a sample. Population is a collection of

all available items as well as each unit

in our study. Sample is a subset of the

population that contains only a few

units of the population. Population data

is used for study when the data pool is

very small and can give all the required

information. Samples are collected

randomly and represent the entire

population in the best possible way.

Measures of central tendency. The

central tendency is a single value that

aids in the description of the data by

determining its center position.

Measures of central tendency are

sometimes known as summary statistics or

measures of central location. The most

popular measurements of central tendency

are mean, median, and mode. The normal

distribution is a bell-shaped

symmetrical distribution in which mean,

median, and mode all are equal. The

curve over here shows the bell-shaped

curve or the normal distribution of

variable X. The point over here that is

X1 is the point which represents the

mean, median and mode of this

distribution. Mean mean is calculated by

dividing these sum of all data values by

the total number of data values. It gets

affected when there are unusual or

extreme values. It is sensitive to the

outliers. Mean can be calculated as

summation over all the values of X in a

collection divided by the size of the

collection.

For example, we have a collection where

we have values as 7 3 4 1 6 and 7.

We find out the sum of these values

which is 28 and there are total of six

values. So 28 / 6 gives us a mean value

of 4.66.

Median,

it is the middle value in the set of the

data that has been sorted in ascending

order.

It is a better alternative to mean since

it is less impacted by outliers and

skewess.

It is closer to the actual central

value.

Median is calculated differently for

different sizes of data.

Differentiated as if the total number of

values is odd or if the total number of

values is even. If the size of the data

is odd. For example, in this case we

have five elements.

After sorting whatever middle value we

get

that means n + 1 by 2 term in this case

5 + 1 / 2

that is the third term which is four is

the median value.

In case when the total number of values

is even like here there are six values.

The average or the mean of the two

central values is considered as the

median. In this case the median is the

mean of 6 and four which is five. Mode.

Mode represents the most common value in

the data set. It is not at all affected

by extreme observations.

It is the best measure of central

tendency for highly skewed or non-normal

distribution.

Mode for categorical data is determined

by estimating the frequencies for each

categories

and then the category with the highest

frequency is considered to be mode.

Like in this case 7 has the highest

frequency. Hence seven becomes the mode

value. However, in case of continuous

data or quantitative data, the

calculation of mode is slightly

different. The first step in calculation

of mode is dividing the data into

classes which are equal with then

getting the frequency of data points

lying in within that range of classes

and finally selecting the class with the

highest frequency.

Using the range of that class and the

frequencies, we can get the final mode

value.

Using the formula L+

minus F_sub_1 multiplied to H divided by

FM minus F_sub_1 plus FM minus F_sub_2.

Here L is the lower limit or the lower

observation of the mode class.

H is the size of the mode class.

FM is the frequency of the mode class.

F_sub_1 is the frequency of the class

proceeding to mode. And F_sub_2 is the

frequency of the class succeeding to

mode. This gives us the final mode

value.

Mean versus expectation.

Now let's talk about mean versus

expectation.

So in general we use the expected value

or expectation when we want to calculate

the mean of a probability distribution

that represents the average value we

expect to occur before collecting any

data. And mean on the other hand mean is

basically used when we want to calculate

the average value of a given sample.

This represents the average value of raw

data that we may have already collected.

We can understand this by using a simple

example.

Now to calculate the expected value of

this probability distribution, we can

use a specific formula from the previous

discussion.

This is going to be the expected value

where X is going to be the data value

and this PX is the probability of value.

For example, we could calculate the

expected value for this probability

distribution to be as shown.

So here it will be 1.45 goals.

So this represents the expected number

of goals that the team will score in any

given game.

And then if you talk about calculating

mean, so we typically calculate the mean

after we have actually collected raw

data.

For example, suppose we record the

number of goals that a soccer team will

score in 15 different games.

Now to calculate the mean number of

goals scored per game,

we can use the following formula

where sum of x is basically the sum of

all the goals divided by n and the

number of records or we can say the

sample size.

It is as shown on the screen.

So this represents the mean number of

goals scored per game by the team.

Measures of asymmetry.

The difference between the three

distinct curves can be studied in this

image.

The central curve is the normal or no

skeus curve. Here mean, median and mode

all lie on the same point. This normal

curve is symmetrical about its mean,

median and mode.

That means the left hand side of the

curve is a mirror image of the right

hand side of the curve.

However, in case of negatively skewed

data, the tail is elongated on the left

hand side

and the mean is smaller than the mode

and the median values or is on the left

hand side of the mode.

Hence indicating that the outliers are

in the negative direction.

On the other hand, in case of positively

skewed, the data is concentrated on the

left hand side of the curve.

While the tail is elongated or longer on

the right hand side of the curve,

the mean is greater than the mode and

median

or is on the right hand side of the mode

and median indicating that the outliers

are in the positive direction.

Let's consider an example.

The graph here shows the global income

distribution for the year 2003 2013 and

a projection for 2035.

If we see the global income distribution

statistics for 2003 it is highly right

skewed.

We can observe in the previous graph

that in 2003

the mean of $3,451

was higher than the median of $1090.

The global income is definitely not

evenly distributed. The majority of

people make less than $2,000 each year.

while only a small percentage of the

population earns more than $14,000.

Measures of variability.

Measures of variability.

Dispersion. The measure of central

tendencies provide a single value that

addresses the full worth. However, the

central tendency cannot depict the

viewpoint entirely. The metric of

dispersion helps us focus on the

inconsistency in the data spread.

Measures of dispersion describe the

spread of the data.

The range, intercortile range, standard

deviation and variance are examples of

dispersion measures.

Range.

The range of distribution is the

difference between the largest and the

smallest amount of data.

The range, for example, does not include

all of a series positive aspects.

It concentrates on the most shocking

aspects and ignores that aren't

considered critical. For example, for a

set 13, 33, 45, 67, 70.

The range is 57. That is the maximum of

this which is 70 minus the minimum over

here which is 13.

Variance.

Variance is the average of all squared

deviations.

It is defined as the sum of squared

distance between each point and the mean

or the dispersion around the mean.

The standard deviation is used as

variance suffers from a unit difference.

Variance can be computed as sigma square

summation over x - mu^ 2

divided by n

where mu is the mean of the data, x is

the individual data point

and n is the size of the data.

This representation is for a population

data.

for a sample data variance can be

computed as X minus

Xar whole square summation

over it divided by n minus one.

Here Xbar is the mean of these sample

data and n is the sample size.

The units of values and variance are not

equal.

So another variability measure is used.

Standard deviation.

Standard deviation is a statistical term

used to measure the amount of

variability or dispersion around a mean.

The standard deviation is calculated as

the square root of variance. It depicts

the concentration of the data around the

mean of the data set.

Standard deviation as indicated

previously can be computed as square

root of variance

for a population data. Standard

deviation sigma can be computed as

square root of summation over x i minus

mu^ square / n

where mu is the mean of the data x i are

the data points and n is the size. Let's

consider an example.

Let's find out the mean, variance, and

standard deviation for this data. The

data values are 3, 5, 6, 9, and 10. To

find out the mean, we first find the sum

of all these data values

that is 33 and divide it by the count,

which is five.

We get the mean of 6.6. To compute the

variance, we start by computing the

deviation.

That is X minus the mean of X. Here 3 is

one of the values of the data and 6.6 is

the mean.

So 3 - 6.6 squared and we do that

to find out sum of all the deviations

divided by the count

which is five.

We end up getting an overall variance of

6.64.

Standard deviation as we know is

measured at square root of variance that

is square of 6.64

which amounts to 2.576.

Measures of relationship.

Measures of relationship coariance.

Covariance is the measure of joint

variability of two variables.

It measures the direction of the

relationship between the variables. It

determines if one variable will cause

the other to alter in the same way.

Coariance between variable X and Y can

be computed as summation over the

product of X I - XR

and Y I - Y bar the whole divided by N

minus one.

Here Xar and Y bar are the mean of X and

Y respectively. The value of covariance

can range from minus infinity to a plus

infinity.

Correlation. Correlation is normalized

coariance.

It measures the strength of association

between two variables. The most common

measure for correlation is the Pearson

correlation coefficient.

Correlation between two variables

X and Y can be measured with respect to

coariance as coariance between X

and Y divided by the standard deviation

of X and standard deviation of Y.

The value of correlation ranges from a

negative 1 to positive 1.

Types of correlation.

Correlation can be either a positive

correlation,

zero correlation or a negative

correlation.

The first picture over here represents a

perfect positive correlation

wherein a straight line with a positive

slope

is representing the relationship between

the two variables.

Zero correlation means that the line

representing the relationship between

the two variables is horizontal to the

xaxis.

Perfect negative correlation can be

represented by a straight line with a

negative slope.

Correlation equals to 1 implies a

positive relationship. That is when one

variable increases the other variable

also increases. A correlation value of

negative one implies a negative

relationship. That is when one variable

increases the other decreases.

The correlation coefficient of zero

shows that the variables are completely

independent of each other.

Let's consider an example.

Here we have two variables height and

weight.

To compute the correlation between

height and weight,

we use the correlation formula as

covariance of X

and Y divided by standard deviation of X

and standard deviation of Y.

Here height is the X variable and weight

is the Y variable.

First to compute coariance we compute

the x - xar and y - y bar values and

then the product of them.

We then compute x - xrยฒ

and y - y bar square values to compute

the standard deviations of height and

weight respectively. Correlation as we

know has been defined as covariance of X

and I and Y divided by standard

deviations of X and Y.

This can also be represented as

summation over x - xr multiplied to y -

y bar

divided by square root of summation over

sum of squared deviations that is x - xr

square multiplied to square root of

summation over y - y bar square that is

sum of square deviations for y.

Now let's find out values to put into

this formula.

First we find out the overall sum of

height to get the mean of height which

is 5.14.

Similarly we get the sum of weight to

get the mean of weight as 50. We now get

the summation over x - xr multiplied to

y - y bar to get the numerator for the

formula. Then we compute x - xr square

summation

and y - y bar square that is sum of

squared deviation of x and y

respectively.

Now we put in the values in this final

correlation formula to get a correlation

value of 0.889.

This indicates that height and weight

have a positive relationship.

It is evident that as height grows,

weight also increases.

In this module, we will be talking about

expectation and variance.

So the expected value or we can say mean

of a given variable that we can denote

by X is a discrete random variable where

it is a weighted average of the possible

values that X can take and each value is

going to be according to the probability

of that specific event occurring.

So usually the expected value of X is

denoted by a simple formula where we can

define the expectation based on the X

parameter

which is going to be the sum of each

possible outcome multiplied by the

probability of the outcome occurring.

So in more concrete terms, the

expectation is what we would expect the

outcome of an experiment to be on

average.

We can take an example for the coin. If

a coin is being tossed 10 times, then

one is most likely to get five heads and

five tails.

Same logic can be discussed if we talk

about another example of rolling a

dieice. So there are six possible

outcomes when you roll a dieice. 1 2 3 4

5 6. And each of these has a probability

of 1x 6 of occurring. So we can say that

the expectation is going to be 1

multiplied by the probability of that

happening which is going to be 1x 6 + 2x

6 + 3x 6 + 4x 6 + 5x 6 + 6x 6 and that

is going to give us 3.5 as an output.

The expected value is 3.5.

So if you think about it, 3.5 is halfway

between the possible values that I can

take and this is what we should have

expected.

Next we talk about the concept of

variance. So variance of a random

variable allows us to know something

about the spread of the possible values

of the variable.

So for a discrete random variable X, the

variances of X is going to be denoted by

using a simple formula that is going to

be var=

E X - M the whole square where M is

basically the expected value of the

expectation of X. So this is more like a

standard deviation of X which can also

be represented by using this formula. So

the variance does not behave in the same

way as expectation when we multiply and

add constants to random variables.

So now there are two different type of

variance that we can have a fair

understanding on. First of all we have

low variance and then we have high

variance.

So low variance simply means that there

is a small variation in the production

of the target function with changes in

the trading data set and at the same

time high variance as we can see here

high variance shows a large variation in

prediction of the target function with

changes in the trading data set. So a

model that shows high variance learns a

lot and perform well with the training

data set and it does not generalize well

with the unseen data set and that's why

as a result such a model gives good

results with training data set but shows

high error rates on the test data set

and since the high variance a model

learns too much from the data set it

leads to an overfitting of the model. So

model with high variance will be having

couple of issues like it may lead to

overfitting or it may also lead to

increase in model complexities.

Next we have skewess.

So skewess in simple terms is basically

a measure of asymmetry of a

distribution. So distribution is

asymmetrical when its left and right

sides are not the mirror images.

Right now this is a mirrored image and a

distribution can have right positive or

we can say negative or it can have zero

skewess.

So right skewed in this scenario is

basically the distribution is longer on

the right side of its peak

and a left skew distribution is going to

be we can say where it is longer on the

left side.

So we can see we have this one as a part

of right side. It is more elongated

towards the right side and this one is

more elongated towards the left side. So

we can think of skewess in terms of

tails. A tail is long tampering and the

end of a distribution. So it simply

indicates that they are observations at

one end of the distribution but that

they are relatively infrequent. So a

right skew distribution has a long tail

on the right side as you can see here.

So the number supports observed. Let's

say we have a data on a per year basis.

So again we can have a more skewess

towards the right side where data is

being dropping as we continue to

increase the number of years. For

example we may have a high sales towards

the beginning of year suppose in 2022

but again as we proceed to 2023 second

half we are seeing the dip in

performance. So that is rightly skewed

and same way let's suppose if we started

with the sales figure it was really less

in suppose 2002

but again as we proceeded to 2023 now

our sales have been gradually

increasing. So it's more like skew

towards the left section as a part of

negative skew. Next we have curtosis.

So curtosis is basically a measure of

the tailness of a distribution.

So taeness is how often the outliers

occur and act as curtis is the tailness

of the distribution related to a normal

distribution. So a distribution with

medium curttosis is called as meocurtic.

A distribution with low curtosis like

this one. This is called as the

platicurtic and then distribution with

high curtosis like this one. This is

called as the leptoccuric.

So tails here they are tapering ends on

either side of a distribution like this.

So they represent the probability or the

frequency of values that are extremely

high or extremely low to the mean.

In other words, tails here represents

how often the outliers occur.

So there are three type of curtis. We

have platocurtic which is negative,

leptocortic which is a positive towards

the upper end and then we have messertic

which is a normal distribution. So

messertic is the medium tail. So normal

distributions they have a curtosis of

three. So any distribution with a curtis

of a prox value of three is going to be

messertic. And curtosis is described in

terms of excess curttosis which is

curtosis minus3. And since normal

distribution they have a curtosis of

three axis curtises makes comparing a

distribution curtosis to a normal

distribution even easier. Introduction

to probability.

Probability theory. Probability is a

measure of the likelihood that an event

will occur.

Let's consider an example of coin toss

where the chances of getting heads on a

coin are 1 by two or 50%.

The probability of each given event is

between zero and one both inclusive.

Sum of an events cumulative probability

cannot be greater than one.

Hence the probability of an event x lies

between zero and one. This means that

the integral of probability of

distribution over x equals to 1.

Conditional probability. Conditional

probability of any event A is defined as

the probability of occurrence of A given

that event B has previously occurred.

Condition probability of event A given B

can be estimated as probability of A

intersection B that is probability of

both A and B happening together

divided by the probability of B.

It is also written as that probability

of A intersection B equals to

probability of A given B multiplied to

probability of B.

Let's consider an example.

In a coin, we are doing a two coin flip.

Coin one gets heads, tails, heads, and

tails in subsequent flips.

while coin 2 gets tails, heads, heads,

and tails in the subsequent flips. Now,

the probability that coin one will get a

head is 2 out of four. While the

probability that coin two will get heads

is again two out of four.

The probability that both coin one and

coin two will have a heads is just one

out of the four flips.

Hence the probability that coin one will

get heads given that coin 2 is already

heads can be computed as probability of

coin one edge intersection coin 2 edge

that is 1x4 divided by probability of

coin 2 edge

that's a given that is 2x 4 which is

going to be 0.5 or 50% based

base theorem Base theorem calculates the

conditional probability of an event

based on its prior probabilities.

Basically base theorem incorporates the

prior probability distribution to

predict the posterior probabilities base

theorem for conditional probability

can be expressed as probability of A

given B equals probability of B given A

divided by probability of B multiplied

to probability of A.

Base theorem allows updating the

probability values by using new

information or evidence. Here

probability of A is known as prior

probability. That is the probability of

event that before any new data is

collected. Probability of A given B is

known as the posterior probability. It

is the revised probability of an event

occurring after taking into

consideration the new information

probability of B given A is known as the

likelihood and probability of B is

probability of observing an evidence B

model. An example consider an example

for calculating the likelihood of having

diabetes based on frequency of fast food

consumption. Here is the observed data.

Let's say the fast food audience is 20%.

Diabetes prevalence is 10% and 5% is

fast food and diabetes.

The chances of diabetes given fast food

that is the conditional probability of D

given B can be calculated as probability

of diabetes and fast food together

divided by probability of fast food.

That means 5% divided by 20%. that

equals 25%.

Define an analysis can state eating fast

food increases the chance of having

diabetes by 25%.

The multiplication rule of probability

if events A and B are statistically

independent and probability of A

intersection B can be given as

probability of A given B multiplied to

probability of B. However, probability

of A intersection B is also given as

probability of A multiplied to

probability of B. Here probability of A

given B equals to probability of A when

we assume that probability of B is non

zero. Similarly, probability of B equals

probability of B given A assuming

probability of A is non zero. Chain rule

of probability joint probability

distributions over many random variables

can be reduced into conditional

distributions over a single variable. It

can be expressed as probability of X1 X2

so on until XN equals probability of X1

intersection probability of X I given

probability of X1 till X I minus one.

For example, the joint probability of A,

B and C can be given as probability of A

given B. C multiplied to probability of

B given C multiply to probability of C.

Logistic sigmoid.

The logistics function is a type of

sigmoid function that aims to predict

the class to which a particular sample

belongs. Its outcome is discrete binary

value. a probability between zero and

one. The logistic sigmoid is a useful

function that follows the yes curve. It

saturates when the input is very large

or very small. Logistic sigmoid is

expressed as sigma of x= 1 upon 1 + e to

the power minus x.

The logistic sigmoid can be expressed as

sigmoid function of x is given as 1 upon

1 + e ^ minus x where e is the ooler's

number.

Gshian distribution.

The gossian distribution is a type of

distribution in which data tends to

cluster around a central value with

little or no bias to the left or right.

It is often referred to as normal

distribution.

In absence of prior information, the

normal distribution is frequently a fair

assumption in machine learning

equation.

The formula for calculating Gaussian

distribution is described as the normal

distribution of X.

That is the function of x given mean as

mu and variance is sigma square can be

calculated as 1 upon sigma square

roo of 2 pi e to the power -/ x -

mood / sigma square

where mu is the mean or peak value which

also is the expected value of x.

Sigma is the standard deviation. Sigma

square is the variance.

A standard normal distribution has a

mean of zero and a standard deviation of

one.

Goshan distribution can be univariate

which describes the distribution of a

single variable X.

It can also be multivariat where it can

just use to describe the distribution of

several variables.

It is represented in 3D of ND formats.

Law of large numbers.

Now let's talk about law of large

numbers. The law of large numbers states

that an observed sample average from a

large sample will be close to the true

population average and that it will get

closer in the larger sample. So the law

of large number does not guarantee that

a given sample spatially a small sample

will reflect the true population

characteristics or that a sample does

not reflect the true population will be

balanced by a subsequent sample. This is

for the law of large numbers to express

the relationship between scale and

growth rate.

So there are multiple examples through

which we can understand

and it is widely used in statistical

analysis in working with the central

limit theorem in terms of the business

growth. So there are multiple real time

setup in which these are going to be

used. So if you talk about tossing a

coin so tossing a coin in a number of

times will give us two different type of

outcomes.

the result will spread evenly between

head and tails and the expected average

value is going to be half.

That means 50 times tails and 30 times

heads. But again, if you toss a coin

1,000 times, then the result can be in

different manners because out of 1,000,

let's say 850 times it has been head and

only 150 times it has been tails and so

on. So that's why the possibility of one

event occurring is going to be changed

in large sample sets as compared to a

small sample sets as in let's say 10

times. So the number of heads and tails

unbalanced for lower number of trials.

So we can see it is unbalanced.

But again as soon as we toss more number

of coins more leans towards the balance

value or we can see the observed

averages.

Next we have p value.

So p value is basically a number

calculated from the statistical test

that describes how likely we are to have

found a particular set of observations

if the null hypothesis were true. So p

values are used in hypothesis testing to

help decide whether to reject the null

hypothesis.

And the smaller the p value, the more

likely we are to reject the null

hypothesis.

So we have a term called as null

hypothesis. So all statistical tests

they have null hypothesis. So for most

tests the null hypothesis is that there

is no relationship between our variables

of in first or that there is no

difference among groups. For example in

a two-taile t test the non-hypothesis is

that the difference between two groups

is going to be zero.

So p value is going to tell us how

likely it is that our data could have

occurred under the null hypothesis.

It is done by calculating the likelihood

of a test statistic

which is the number calculated by a

statistical test using our data. So p

value tell us how often we would expect

to see a test statistic as extreme or

more extreme

than one calculated by a statistical

test. if the null hypothesis of the test

was true.

So there are multiple limitations as

well. So first one is the results can be

significant but again they are they may

not be practical as we have compared it

can be based on multiple hypothesis for

a game for the healthcare test. If the

test is going to be positive or not it

may show even values of the effect of a

variable but not the magnitude in real

life. What exactly is going to be the

application of a drug test being failed

in pharma company? Therefore, it is

recommended to use confidence and levels

in addition to the p values to quantify

or we can say to give a solid figure to

the reserve which we are going to get.

The p values they are interpreted as

supporting or we can say refuting the

alternative hypothesis.

So p value can only tell you whether or

not the null hypothesis is supported. It

cannot tell us whether our alternative

hypothesis is true or why. So the risk

of rejecting the null hypothesis is

often higher than the p value. So

especially when we are looking at a

single study or when using small sample

sizes. So this is because the smaller

frame of reference, the greater are the

chance that as we stumble across a

statistically significant pattern

completely by accident.

Key takeaways.

Key takeaways. Probability and

statistics structure the premise of the

data. The data helps in anticipating the

future or gauging in view of the past

patterns of information.

The central tendency is a single value

that helps to describe the data by

identifying these central positions. The

mean, median and mode are the measures

of central tendencies.

The distribution where the data tends to

be around a central value with a lack of

bias or minimal bias towards the left or

right is called as gshian distribution.

So now let's dive into the definition of

the probability distribution function.

What is probability distribution

function? A function which defines the

relationship between a random variable

and its probability such that you can

find the probability of the variable

using the function is called a

probability density function.

In simple words, probability density is

the relationship between an observation

and the probability. Some outcomes of a

random variable will have low

probability density and other outcomes

will have a very high probability

density. Basically, the probability of a

variable X happening or occurring will

vary and it can sometimes take on a

lower value or it can take on a way

higher value.

The overall shape of the probability

density is referred to as probability

distribution. And the calculation of

probabilities for specific outcomes of a

random variable is performed by a

probability density function or PDF for

short. Now consider a variable X with a

PDF of f ofx.

This is what your probability density

function will look like. There might be

a point where the probability of X

occurring is very high. Hence your

probability distribution function or f

ofx will also be very high. At other

points the distribution or the

probability of X happening or occurring

is going to be very low. Hence your f

ofx is also going to have a very small

value. Basically given the random sample

of a variable we might want to know

things like the shape of the probability

distribution. This here is something

called a normal distribution where a

probability distribution function takes

on a bell shape.

However, this is not the probability

density function that might always

occur. There are different probability

distribution functions and all of their

graphs look very different from each

other. Knowing the probability

distribution for a random variable can

help you calculate movements of the

distribution like the mean and variance.

But it can also be useful for other more

general considerations like determining

whether an observation is unlikely or

very unlikely and might be an outlier or

an anomaly like consider this graph

itself. In this graph, these points over

here which have very less probability

distribution

are outliers which means that the chance

of them occurring is very low. And

basically this is not something that

you're going to see in your regular

scenario for your variable X. Now let's

consider two points A and B which are

values that a variable X can take. P of

A and P of B just represent the

probability of A and the probability of

B which can be found out by drawing a

straight line and coinciding it with our

graphs. The area under the graph over

here which is going to give you your

probability of this region occurring can

be written as probability of A less than

equal to X which is a probability that

we're searching for here less than equal

to probability of B. What does this mean

exactly? This means that this area is

always going to be greater than or equal

to the probability of A but less than or

equal to the probability of B. This

gives us the narrow region

of the probability which is present over

here. And doing this we can find the

probability of occurrence for any value

of X. Suppose you want to find the

probability of B happening. For a

probability distribution function, the

probability of B happening is not simply

this point here, but the entire area of

the graph which is taking place before

this point itself. So if you want to

find the probability between these

regions, you're going to have to find

the entire area and not simply the

probability at one point.

Now, so far we've been talking about

different types of variables which is

discrete random variables and continuous

variables. What exactly do these mean? A

variable which can only take a value

within a certain range is called a

discrete random variable. The value is

usually within a certain distance of

another finite value. An example of this

would be the sum of two dices. Basically

values which are well defined are called

discrete values or and a variable which

has well- definfined values will be

called a discrete random variable.

This variable can only take values which

fall within a certain set of values.

Let's say you roll a dice. The dice can

only give you specific outcomes which

range from 1 to six. This is what you

would call a discrete output.

On the other hand, a continuous random

variable can take on infinite different

values within a range of values. For

example, the height of a student. The

height of a student is not fixed. Even

if the height is 1.7 m, in reality, the

height can be 1.77 or 1.765

or 1.789.

The exact height is very hard to

determine because it's not easy for us

to find the precise value of the height

of a student. So basically the height

can take on an infinite different range

of values. When we're trying to define

the values that a continuous random

variable can take, we usually say it in

the form of a range of values which

means that the value can fall in that

range and can take on any value in that

range. It's not like a discrete random

variable where you can define definitive

values.

Now let's understand a probability

density function with the help of a

graph. Consider the graph below which

shows the rainfall distribution in an

year in a city. The x-axis has the

rainfall in inches or the amount of rain

that we're getting and the y-axis has

the probability density function of

getting that amount of rain.

The probability of some amount of

rainfall is obtained by finding the area

of the curve to the left of it. So let's

say we have a 3.

If you want to find the probability of 3

in of rainfall occurring, we would have

to find the area of the curve

which falls to the left of three. When

we draw a line from three which

intercepts the graph and further extend

it onto the yaxis, we get a value of

0.5.

Simply put, this means that the

probability of 3 in of rainfall

occurring is going to be lesser than or

equal to 0.5. The exact probability can

be found out by finding the area of the

curve

which falls to the left of three.

How do we find the probability

distribution function?

The first step is to summarize your

density with the help of a histogram.

The first step in a density estimation

is to create a histogram of the

observations

in the random sample.

Now what is a histogram? A histogram is

a plot which involves first grouping the

observation into bins and counting the

number of events that fall in each bin.

The counts or frequency of observation

in each bins are then plotted as a bar

graph with the bins on the x-axis and

the frequency on the y-axis. The choice

of the number of bins is important as it

controls the coarseness of the

distribution and in turn how well the

density of the observation is plotted.

It is a good idea to experiment with

different bin sizes for a given data

sample to get multiple perspectives or

views on the same data.

At the same time, the number of bins is

important as it determines how many bars

the histogram will have and their

widths. This will change not only the

shape of the graph but also how the

graph is read. This will also determine

how our density is plotted. Now let's

see how we can summarize our density

with histograms using Python. First

let's import all of our necessary

modules which we're going to require.

We're going to require Mattplot lib to

plot graphs. We're going to need the

normal random function so that we can

get a normal distribution. We're going

to import mean and standard deviation

from numpy to use on our graphs and also

going to normalize our uh data. So we're

going to import the nom function from

sci.

We finished importing all of our

necessary modules. Now let's generate a

sample

which has a size of thousand and it's

going to be a normal distribution. And

we're going to also plot this with the

help of a histogram in bins of 10.

So as you can see here you get a normal

distribution which is nothing but a

almost bell-shaped curve and we have 10

bins here which are centered at zero and

which extend from minus3 to 3.

How will our graph look if we change the

number of bins though?

Let's run it and see. So you still have

a normal distribution but it's not as

well defined because of how less the

number of bins are. you lose a majority

of the data which will contribute to

your normal distribution. It doesn't

look like a proper normal distribution

but it looks more like a discrete data

at this point. Now let's take a look at

the next step of finding a probability

distribution function.

The next step is called parametric

density estimation. What exactly is

parametric density estimation? The

probability density function is of many

types. The shape of your histogram will

help you determine what type of a

function it is. We can also calculate

the parameters associated with the

function to get our density. Now

different probability distribution

functions will have

different graphs which will have

different shapes and which will also

have different parameters like mean,

standard deviation etc associated with

them. Using these parameters, we can

find important points of our data.

Hence, it's very important for us to

recognize what type of a distribution it

is. Common distributions will occur

again and again in different and

sometimes unexpected domains.

Getting familiar with common probability

distributions will help you identify a

distribution from a histogram. And once

identified, you can attempt to estimate

the density of the random variable with

a chosen probability distribution. This

can be achieved by estimating the

parameters of the distribution from a

random sample of data. Now, an example

of this would be a normal distribution

which has two main parameters, the mean

and standard deviation. Given these two

parameters, we will now know the

probability distribution function. These

parameters can be estimated from data by

calculating the sample mean and sample

standard deviation. This entire process

is known as parametric density

estimation and it includes identifying

your probability distribution function

and getting the parameters which are

associated with it.

Now once we have estimated the density,

we can check if it's a good fit.

This can be done in three different

ways. One is plotting the density of the

function and comparing the shape to the

histogram. The next is sampling the

density function and comparing the

generated sample to the real sample. And

the last one is using a statistical test

to confirm if the data fits the

distribution. Now over here as you can

see all we've done is taken our data and

plotted the density function on top of

our histogram and we've compared the

shape. So the distribution so the

density function that we're actually

considering here is a normal

distribution and from this graph we can

see that it's almost an exact fit to our

histogram. Now let's see how we can

perform parametric density estimation

using Python. To begin with,

let's generate a random sample of

thousand observations from a normal

distribution with a mean of 50 which is

determined by the LOC parameter and a

standard deviation of five which is

determined by the scale parameter.

Now just to show you what the

distribution looks like, we're going to

plot it in the form of a histogram. So

this is what the histogram looks like.

But this is just to give you a basic

idea of our data and what it looks like

once plotted. But let's assume that we

don't know the probability distribution

and and we don't know what it looks like

as a histogram and we don't know that

that it's normal. So now if we just

assume that it's normal, we can

calculate the parameters of the

distribution specifically the mean and

the standard deviation.

We would not expect the mean and

standard deviation to be 50 and five.

Exactly given the small sample size and

the noise in the sampling data.

So because of this noise and the small

sample size, we have a mean of almost 50

and a standard deviation of a little

more than five.

Now let's define the distribution as

normal. So now using this we've defined

a normal distribution. We've used the

norm method of the sci-fi uh library and

uh we're doing this with the mean and

the standard deviation that we've

obtained from our samples. So up until

now we're just assuming that it's a

normal distribution and because of that

the parameters that we've calculated is

the mean and standard deviation

and using the calculated mean and

standard deviation we've gotten a normal

distribution.

And up until now again keep in mind we

do not know for certain that it is a

normal distribution. So far all we have

is this data.

So the next thing that we're going to do

is fit the distribution with these

parameters

and then sample the probabilities for a

distribution for a range of values in

our domain which in this case is 30 and

70. So all we're doing is we're

calculating probabilities for a range of

outcomes. And in this case we've taken

30 and 70 as our domain.

So these are the probability

distribution values for the normal

distribution that we've defined over

here. And this is going to uh this is

basically going to give you the outline

of your normal distribution.

These are the points at which your

normal distribution will be plotted. Uh

now what we're basically going to do is

we're going to plot our histograms using

the samples that we've already generated

along with the values and probabilities

of the normal function that we defined

over here.

So as you can see it's an all it's

almost a complete fit. The normal

distribution that we have here is made

using the mean and the standard

deviation of our actual samples.

The reason we took mean and standard

deviation was because we assumed it was

a normal distribution and the parameters

associated with the normal distribution

are mean and standard deviation. Using

the mean and standard deviation, we got

the normal distribution. We calculated

probabilities for this normal

distribution using a random domain of 30

and 70

and we plotted the probabilities and the

values on top of our histogram to see if

the normal distribution was a fit to our

histogram. If it was not a fit, you

would have to go and do the same

procedure with other common probability

density functions

until you found a function which was a

proper fit to your histogram. Now let's

move on to the final step which is used

in the calculation of a PDF. This final

step is called nonparametric density

estimation and it's only used when the

shape of a histogram doesn't match a

common probability density function or

it cannot be made to fit one. In this

case, we will calculate the density

using all samples in our data using

certain algorithms.

This is only done when a data sample

does not resemble a common probability

distribution or it cannot be easily made

to fit the distribution. And this is

often the case when the data has two

peaks. This is also called a biodal

distribution or it has many peaks which

is also called a multimodal

distribution. In this case, the

parametric density estimation will not

be feasible and alternative methods can

be used that do not use a common

distribution. Instead, you will use an

algorithm which is used to approximate

the probability distribution of the data

without a predefined distribution which

is also referred to as a non-parametric

method because we're not using any

predefined parameters. The distribution

will still have parameters but these are

not controllable in the same way as a

simple probability distribution. For

example, a non-parametric method might

estimate the density using all

observations in a random sample in

effect making all observations in the

sample parameters.

Now consider this graph which has two

peaks. You this is not a normal

distribution or any other sort of

distribution that we are familiar with.

So for this we're not going to use a

parametric estimation method but we're

just going to calculate the parameters

for every single sample point in this.

Perhaps the most common nonparametric

approach for estimating the probability

density function of a continuous random

variable is called kernel smoothing or

kernel density estimation or KDE for

short. Kernal density estimation is a

nonparametric method for using a data

set to estimate probabilities for new

points.

It uses a mathematical function and

smoothing probabilities. So the so the

sum of the resultant probabilities is

always one. Now in this case a kernel is

a mathematical function that returns a

probability for a given value of a

random variable. The kernel effectively

smooths or interpolates the

probabilities across a range of outcomes

for a random variable such that the sum

of probabilities always equals one. A

requirement of well- behaved

probabilities. You also have a parameter

called the smoothing parameter which

controls the scope or the window of

observations from the data samples that

contributes to estimating the

probability for a given sample. As such

the kernel density estimation is s is

sometimes referred to as your parsen

rosenbalt window. Now at the end you

also have a basis function which is a

function which is chosen to control the

contribution of samples in the data set

towards estimating the probability of a

new point. This is only done to make

sure that you're not learning from a lot

of noise and that you're not using a lot

of the outliers. Again let's see how we

can perform non-parametric density

estimation with the help of Python.

So first we'll start by importing all

the necessary modules along with the

kernel density estimation which can be

imported from skarn.

Now let's create a biodial distribution

by combining two different samples.

Sample one and sample two. Sample one

has 300 examples with a mean of 20 and a

standard deviation of five. While sample

two has 700 examples with a mean of 40

and a standard deviation of five.

We're then going to use it stack to com

to merge both of them together to get a

final sample.

The means that we've chosen which is 20

and 40 are chosen close together to

ensure that the distributions overlap in

the combined sample.

So this is what our distribution is.

Let's just plot it so you get a basic

idea of what our graph looks like.

So this is what our graph looks like.

Now we already know that none of the

various different uh uh probability

distribution functions fit these graphs.

So now we're going to perform

nonparametric estimations.

To perform nonparametric estimations,

we're going to use the scikitlearn

machine learning library which provides

the kernel density class that implements

kernel density sorry that implements

kernel density estimation. First the

class is constructed with the desired

bandwidth or window size of two

and your basis function

which in this case is a gshian function.

It's a good idea to at least test

different configurations to your data.

And in this case we're only going to try

a bandwidth of two and a gshian kernel.

Uh but usually there are multiple

different kernels that you can uh you

know like uh that you can play around

with and you can also tweak your

bandwidth to exactly fit the

distribution that you have.

Now let's run this.

Uh so now we've gotten our kernel

density estimation. Uh now we can

evaluate how well the density estimates

matches our data by calculating

probabilities

for a range of observations and

comparing shapes to the histogram just

like we did for the parametric case

before. So again we're going to just

calculate different probabilities using

the kernel density function

and we're just going to plot it on top

of a histogram to see how well this the

kernel density function is estimating

for our data.

So these are the probabilities that

we've gotten finally with the con uh

with the kernel density estimation. And

now we're going to plot it on top of our

histogram.

So as you can see it's almost a complete

fit. It's just left out some of these

outlier values which again are ranging

very high. But overall we have a pretty

good fit.

Uh the only problem is it's not very

smooth and you can uh again try tweaking

the bandwidth to different values. Uh so

let's just in this case try tweaking it

to three and see how well it runs. Okay.

So now we've got a new probabilities and

let's run it on top of our bandwidth. Uh

so again we using a bandwidth of three.

You can see that we're fitting our data

even better and we're again

ignoring a lot of the outliers which are

out there. So this is going to give us a

better estimation. The first question

that is probably in your mind is what's

in it for you? What can you expect from

this video?

First we will explain the concept of

regression a machine learning algorithm

to you.

Next we will take a look at the R squar

error which can be used to calculate the

error in regression models.

Next

we will teach you how to calculate the R

squar error and finally we will

implement the R squared error with the

help of Python.

So what is regression?

Regression is nothing but a machine

learning algorithm that helps us

determine the relationship between two

or more variables. It uses input or

independent variables to find the value

of the output or dependent variables.

Regression is a prediction algorithm

which means that given some variables we

can predict the value of an output

variable.

The predicted value is not going to be

from a set of values but it's going to

be a unique value in itself.

Now let's understand what exactly

regression is with the help of a few

independent input variables. In this

case, the variables that we'll be

looking at is rainwater, fertilizer, and

seeds. When we pass these independent

input variables through regression

model, we're going to get a predicted

output.

The output predicted is that a crop will

germinate when all three of these

components are put together in certain

quantities. So with the help of

regression given raw input data we can

find out the dependent output variable

that we'll get

when all of these input variables are

more or less combined. Regression is

nothing but a statistical method which

is used to determine the strength and

character of the relationship between a

dependent variable and a series of other

variables.

Now using regression if you have a set

of data points we can use a regression

model to fit a line which passes through

most of these data points and use it to

predict the outcome for new data points.

The line is fit using equation of a

straight line or a polomial equation.

Now in this graph consider that we have

our dependent variable or the output y

and our independent variable or the

output x. for a certain value of our

independent variable. We are going to

get a certain value of our dependent

variable or our output. Using the data

given here, we can see how our output

varies when our input varies. To predict

the value of our outcome Y, we're going

to need to find a relationship between

all of these data points. To do this,

we're going to plot a straight line

through it. Because, as you can see, all

the data points lie more or less along a

given straight line. Now using the

straight line for any value of x we can

find the approximate value of y. So

suppose you want to find the value of y

at a point x which is given here say

then all you have to do is extend a line

from this point onto a predicted model

which is this line here and then we can

see where this point on the line

coincides with the y-axis and get the

approximate value of the outcome.

The equation of a straight line is given

as y = b + b1 x + e. Where b is a

constant given by the y intercept of our

line or basically where the line

intersects on the y-axis.

B1 is the slope of our line and x is the

point for which we want to find the

output. E in this case is nothing but an

error correcting term.

So this is basically how regression

works and this is how prediction takes

place in a regression model. Next we

will explain the concept of the R squar

error to you. So what is R squar error?

R squared error is nothing but an error

measurement term which calculates how

well a regression model fits the data.

It determines the amount of variance in

a model caused by the input variables.

Now, R squar is a statistical measure

that represents the portion of variance

for a dependent variable that's

explained by an independent variable or

variables in a regression model. R 2 is

used to explain to what extent the

variance of one variable affects the

variance of a second variable.

So if the R square of a model is 0.5

then approximately half of the observed

variation can be explained by the

model's inputs.

In other words, an R squar of 60%

reveals that 60% of our data fits our

regression model exactly.

Now in this case in this graph if we

have a variance of 60% it means that 60%

of our data points fall exactly on our

regression line. However it is not

always the case that a high R squar is

good for a regression model. The quality

of the statistical measure depends on

many factors such as the nature of the

variables employed in the model, the

unit of measure of variables and the

applied data transformation.

Thus, sometimes a high R squar can

indicate the problems with the

regression model. A low R squar figure

is generally a bad sign for predictive

models. Now, all this time we've been

talking about variance in our model.

What exactly is variance? Variance is

nothing but a statistical term which

determines how spread out our data is

and tells us how many outliers are

present in it. Basically, it's a measure

of how far a set of numbers is spread

out from their average value.

Using variance, you can basically figure

out where your data is centered

and how spread it is from the mean and

also you can find out how many outliers

it has.

Now how can you calculate the R squar

error?

Let's start by considering the data that

we have been given as shown below. Now

we can find the relationship between the

input and the output variables by

plotting a straight line or a regression

model that passes through most of the

data. To get the perfect fit for a

model, we don't necessarily need to have

the line passing through as many data

points as possible.

A true measure of a good model is that

we reduce the error which is present in

our model. Now how do you find this

error? The error present in our model is

given by nothing but the distance

between our predicted line and the data

points which do not fall on the line.

This is what we have to minimize.

The distance between our data points and

our line can be calculated by

subtracting

our data point from the point at which

it coincides on our regression line. We

square this just to get rid of any

negative coefficients that may occur due

to finding the difference between the

two points. Now to find the variance in

our data, we're going to find the mean

and subtract the data points from the

mean. This will basically tell us how

spread out our data point is from the

average value. We can then square these

differences and add up the result to get

our total variance. This is also known

as the sum of squares total. And using

this we can find the total variance in

our data. The mean is nothing but the

average of our data. And using the mean

we can find the center of our data.

So this is exactly where our data is

centered. This is the average value that

occurs in our data. Now, the variance is

nothing but the distance of all of our

data points from the mean. If we do

this, we're basically going to find out

how spread apart our data is from the

average value or how far all of our data

points lie from each other.

When we subtract the position of our

data points from our mean and square it

and add all of that up, we get something

known as the sum of squares total.

Now the R squ error is the total

variance in our input data. It can be

obtained by dividing the SSR by the SST

and subtracting the results from one.

So the R squar error totally becomes 1

minus the sum of our squared errors

divided by the sum of the squared

difference between our data points and

the mean. Now this value gives us the

variance and this is why we can say that

R squ is used to find the portion of

variation in our data.

Now how can you implement R squ error

with Python?

To calculate R squared error with

Python, we're going to look at the data

which depicts the weather conditions

which were present during World War II.

And using the variables which are

present in our data, we're going to

create a model which predicts the daily

weather forecast during World War II.

And then we're going to use R squared

error to find the accuracy of our model.

So this is our R square uh model. So

we're going to start off by importing

all of our necessary modules. We're

going to use the model numpy to perform

numerical calculations on our database

and arrays and we're going to use

seaborn and mattplot lib to plot our

data.

So now we've managed to import all of

our data sets.

Let's also load our data by reading in

the CSV file that it is stored as in the

form of a data frame.

So over here as you can see we've read

in the CSV file into uh a variable

called weather and then after that we're

changing weather into a panda's data

frame called climate. This is what our

data frame finally looks like.

So as you can see in our data frame we

have five rows because we're only

looking at the top five rows and we have

31 columns. So these are values which

are not required and which are basically

going to increase our error value. So

let's drop them and get rid of them.

Now let's also drop any

empty values which may occur in the

remaining columns of our data set and

see what the final data set looks like.

So this is our final data set. We have

at the end we're only left with max

temperature, minimum temperature, and

mean temperature.

Now let's plot a count plot of a max

temperature.

A count plot is basically going to go

through the entire max temperature

column and figure out how many times

every temperature value occurs. So it's

going to figure out how many it's going

to count how many times 29.44444

has occurred and it's going to plot that

on this graph and it is going to do that

for every unique temperature value which

is present in our column. So finally

this is the value that we get. Uh so

over here as you can see the majority of

our temperature values are concentrated

within this range. This means that these

temperature values are the ones which

occur most frequently. The other ones

can be considered as outliers because

they rarely are seen in our data

and they can further skew the output

that we're going to get. Now let's do

the same with minimum temperature. Let's

plot a count plot for minimum

temperature.

So for the minimum temperature we can

see a very similar plot to the one that

we got for a maximum temperature. Most

of the values are concentrated around

this region but the outlier values here

are way fewer.

Now let's plot a regression plot between

our maximum and our minimum temperature.

So using the regression plot we can plot

a regression line for the two variables

in our x and y axis. So this is our

x-axis and this is our y-axis. This is

basically going to plot a straight line

which best fits the data

that we are getting here. So over here

as you can see this is a regression

line. This thin blue line is a

regression line which intersects our

minimum temperature at a value which is

between -30 and -40 and it passes

through our entire data. So uh for a

value of maximum temperature which is

zero using this we can predict the

minimum temperature that would have

occurred on the same day.

So for zero it'll be somewhere around -

10ยฐC. So if we saw maximum temperature

of 0ยฐ on that day, we would have seen a

minimum temperature of minus 10 on the

same day. Now let's plot a heat map to

see how these values are correlated with

each other. So the correlation is

basically

used to find which values affect each

other linearly

or which values

when changed will also affect the change

in other values. So over here as you can

see minimum temperature

and maximum temperature have a

correlation of 0.8. 88 which means if

minimum temperature changes then the

maximum temperature will also change to

0.88.

Now the best correlation is obviously

going to be between the mean

temperatures and the minimum and maximum

temperatures. The mean temperature is

nothing but the average temperature

value that we have. So this this is

basically going to lie in the middle of

all of our temperature values which is

why we going to have a better

correlation for mean temperature. But

minimum temperature and maximum

temperature are also pretty well

correlated with a correlation value of

0.88.

This means that if our minimum

temperature fluctuates, our maximum

temperature will also fluctuate

proportionately.

Now let's separate our input and output

values. We're going to predict the

maximum temperature given our minimum

temperature. Here x is our input

variable and y is our output variable.

So now we're basically just going to get

all the important values in our x and y

data sets. So after that this is what

our x and y data sets are going to look

like. Now let's split our data set into

training and testing sets.

The training set will be used to train

our regression model and the testing set

will be used to predict how well our

regression model is performing.

The training data is the data which will

be visible to our model or the data

which a model is allowed to have access

to. Testing data will be data which the

model has never seen before or which it

doesn't have access to and hence it will

be made to work on completely new data

to better test how well we've fitted to

our data set. Now we can split our data

set into training and testing sets by

using the train test split functionality

from our skarn.mmodel selection library.

Now finally from our scikitlearn library

let's import a linear regression model.

We're going to initialize a linear

regression model to a variable called

regressor and then we're going to fit a

linear regression model to our training

data set.

So we finally got our trained linear

regression model. Now let's use this

model to perform predictions on our

testing data set. So these are the

values that we've gotten after running

our linear regression model on our

testing data set. Let's see how well

we've performed.

We are going to import the R2 score from

our skarn metric.

Now the R2 score will directly perform R

squared error on our prediction and

testing data set and see how well our

testing data set matches a prediction

data set.

So now we've gotten an R squared error

of 0.9345

which basically means that 93% of our

output values are influenced by our

input values. This also means that a

model is 93% accurate

and that approximately

93% of our observed variation can be

explained by the model's inputs. Ever

wondered how to build an AI project that

actually gets noticed by Google, OpenAI,

or top startups, not just a chartboard

or recycled homework. Today I'm going to

walk you through 10 AI project ideas for

26 that are practical, futuristic and

portfolio ready. I'll tell you exactly

which models, framework and data sets to

you so that you can start coding

immediately. Now before we jump into hit

that like button, share and subscribe

because keeping up with future proof AI

projects is going to give you a massive

edge. Let's start with the AI shopping

buddy. This project acts like a personal

stylist and interior designer. Users

upload photos of the room, outfit, or

even face, and the AI suggest products

that match color, style, and

preferences. This isn't just about

throwing recommendations at someone.

It's about computer vision to understand

images, generative AI to create style

suggestions, and recommendation

algorithms to find the perfect products.

Personalized recommendation systems

drive massive engagement and conversions

which is why companies like Amazon,

Flipkart, Myntra or Urban Ladder would

be thrilled to hit someone who can build

this. Completing a project like this

demonstrates skills in deep learning,

computer vision, generative AI and full

stack deployment for web or mobile app.

While shopping and lifestyle AI is

exciting, the next project takes up to

our health and wellness. The smart

health analyzer predicts stress burnout

or sleep issues by analyzing voice,

facial microp expressions and variable

data. It uses multimodel AI that

integrates time series analysis for

variable data. NLP for voice and text

and computer vision for micro

expressions. Health tech startups in

India and globally like healthy, cure

fit, Fitbit and Apple Health are looking

for engineers who can make predictive

wellness tools. Building this project

demonstrates your ability to work with

multimodel AI, pre-process complex data

sets, train models, and visualize

result. Moving from personal health to

professional efficiency, the AI

productivity agent automates your daily

workflow. It reads emails, scans your

calendar, understands priorities, and

builds an optimized schedule. It uses

NLP to parse emails, API integration

with other tools like Gmail, Slack, and

Notion, and optimization algorithms to

prioritize task efficiently.

Productivity loss is a major issue for

companies which is why tech giants like

Google Workspace, Microsoft 365 and

startups in workflow automation would be

very interested in this project. It is a

great way to demonstrate automation, NLP

API integration and practical problem

solving skills. Taking automation to the

next level, the voice toaction system

allows users to speak commands and have

the AI perform multi-step action such as

booking flights, organizing files or

generating reports. It relies on

speechtoext models, intent

classification using NLP and task

automation pipelines. You can train

intent classification models using data

set into the snips NLU data set. This

project builds directly on productivity

AI and is exactly the kind of work that

Amazon openai or Apple would notice for

voiced driven automation solutions. Once

we have automated task, why not explore

creativity? Generative AI story maker

allows you to create full stories

including scripts, characters and

visuals based on just a few keywords.

Now it uses large language models for

text generation and image generation

models like stable diffusion deli3 for

visuals and you can also train

fine-tuned models on data sets like CMU

book summary corpus on writing prompts

text to speech library such as scope TTS

or GTTS can add narration media

companies like Netflix, Ubisoft and

Adobe are actively investing in

generative AI and a project like this

would definitely stand out. Building on

the idea of multiple AI capabilities

working together, the multi- aent AI

team project introduces collaboration

between AI agents. Multiple AI agents

are assigned specialized roles such as

researching, writing, criticing, and

summarizing. They communicate and

coordinate to complete complex task

using multi- aent reinforcement,

learning and communication protocols.

Enterprise AI and automation platforms

are investing heavily in this approach

and companies like Enthropic, OpenAI, AI

workflow startups are actively seeking

engineers who can build collaborative AI

systems from collaboration to

observation. The AI body language reader

analyzes micro expressions, tone of

voice and posture to provide feedback on

communication skills. It can be applied

in interviews, public speaking or remote

coaching. Computer vision tools such as

open pose or media pipe pose track

gestures and posture. While audio

processing libraries such as librosa

analyze tone models can be trained using

data sets like Raves for audio and CK

plus for facial expressions. HR tech

companies like High View, Pytrics and AI

coaching startups would highly value

this type of project as it help bridge

human behavior and AI analysis. Nucation

is also another area being transformed

by AI. The personalized tutor with

adaptive difficulty creates an AI tutor

that adjusts lessons in real time based

on student performance. Knowledge

tracing models like deep knowledge

tracing combined with transformers for

content generation allow the AI to adapt

to each learner. Data sets such as

assessments or edn nets can be used for

training. Now this AI can generate new

questions, explanations and motivational

feedback based on learning pace.

Companies like Baiju, Vidanto, Corsera

and Udemy are constantly looking for

talent that can build adaptive learning

platforms. Next, we move into research

augmentation with the autonomous

research agent. This AI can answer

research questions by reading academic

papers, extracting insights, summarizing

information, and citing sources

automatically. It uses the S2 or data

set for academic papers. Cybboard for

scientific text embeddings and hugging

face transformers or lang chain for

reasoning and summarization. Citation

extraction can be done with NLP passers

or rejects academic platforms. AI labs

and companies like Google research, open

AI, research gate and LCV would hire

engineers who can build this system.

This project connects perfectly with

education focused AI extending learning

into automated research capabilities.

Finally, we arrive at realworld robotics

control with the AI. The ultimate

demonstration of cuttingedge skill. This

project trains AI to control robotic

arms or humanoids based on goals rather

than just rigid instructions. It uses pi

bullet or vbots for simulation. Stable

baseline 3 or R lib for reinforcement

learning and open CV or media pipe for

vision input. Sim to real transfer

techniques bring simulations into realw

world scenarios. Robotics companies like

Boston Dynamics, Appronic, Amazon

Robotics and Agibot are seeking

engineers capable of endtoend AIdriven

robotic systems. After exploring

softwarebased AI projects, robotics is

the next step to show mastery of AI

applied in the physical world. These 10

AI projects are more than just ideas.

>> We will learn about some of the machine

learning and deep learning interview

questions.

So let's begin with our first question.

The first question is how to detect

outliers in data. So in data analytics

and machine learning, you often find

data points that lie at an abnormal

distance from other points in a random

sample from a population. Those are

called outliers. Now outliers in data

can significantly impact any prediction

analysis. There are majorly three

different methods to treat outliers.

First we have the univariate method. It

is one of the simplest methods for

detecting outliers. The univariate

method uses box plots. A box plot is a

graphical display for describing the

distributions of the data. Box plots use

the median and the lower and upper

quartiles.

This method looks for data points with

extreme values on one variable. Next, we

have the multivariate method. So, the

multivariate outliers can be found in an

n- dimensional space having n features.

We look for unusual combinations of all

the variables in this method. Finally,

we have Minowski error. This method

reduces the contribution of potential

outliers in the training process. The

Minowski error is a loss index that is

more insensitive to outliers than the

standard mean squared error. Now moving

on to the second question. What is a

confusion matrix? So a confusion matrix

is a table that is used to describe the

performance of a classification model on

a set of test data for which the true

values are already known. The target

variable has two values positive or

negative. The columns represent the

actual values of the target variable

which you can see here. The rows

represent the predicted values of the

target variable which you can see here.

Now there are four important terms that

are related to confusion matrix. First

we have true positive which is this one.

So in true positive the predicted value

matches the actual value. So the actual

value was positive and the model also

predicted a positive value. Then we have

true negative which is also represented

as tn. The true negative depicts the

predicted value matches the actual

value. Now the actual value was negative

and the model predicted a negative

value. Next we have false positive. Now

false positive is also known as a type

one error. In false positive the

predicted value was falsely predicted.

The actual value was negative but the

model predicted a positive value.

Finally we have false negative. A false

negative is also known as type two

error. So in false negative the

predicted value was falsely predicted.

The actual value was positive but the

model predicted a negative value. Now

moving to our third question which is

explain the ROC curve. Now the ROC curve

is one of the most important evaluation

metrics for checking the performance of

any classification model. ROC stands for

receiver operating characteristic.

Receiver operating characteristic or ROC

curve is a method to compare the

diagnostic tests. The ROC curve is

created by plotting the true positive

rate against the false positive rate at

various threshold settings. So here on

the y-axis you have the true positive

rate. On the x-axis we have the false

positive rate. The true positive rate

indicates the proportion of observations

that were correctly predicted to be

positive out of all positive

observations. Similarly, the false

positive rate is the proportion of

observations that are incorrectly

predicted to be positive out of all

negative observations.

You can take an example. Suppose in

medical testing, the true positive rate

is the rate in which people are

correctly identified to test positive

for the disease in question. Let's say

the corona virus testing. ROC does not

depend on any class distribution. This

makes it useful for evaluating

classifiers predicting rare events such

as diseases or disasters. Now moving to

the fourth question we have what are the

assumptions for linear regression. So

linear regression analysis is used for

modeling the relationship between a

single dependent variable Y and one or

more feature or predictor variables.

Some of the important assumptions for

linear regression are so first they

should have linearity. So linear

regression needs the relationship

between the independent and the

dependent variables to be linear. It is

also crucial to check for outliers since

linear regression is sensitive to

outlier effects. Next we have

homocyasticity.

Homoscadasticity

illustrates a situation in which the

error term that is the noise or random

disturbance in the relationship between

the features and the target variable is

the same across all levels of the

dependent variables. Third we have

independence. So observations should be

independent of each other. Finally, we

have no multi-olinearity.

So there should be little or no

multi-olinearity.

Independent variables should not be too

highly correlated. Now moving to our

fifth question in our list of interview

questions.

The question is what is regularization

in machine learning? Explain the L2

regularization.

So regularization is a machine learning

technique that is used to reduce the

errors by fitting the function

appropriately on the training set in

order to avoid overfitting of data. So

overfitting happens when a model learns

the detail and noise in the training

data to the extent that it negatively

impacts the performance of the model on

new data. So here you can see we have a

nice plot which shows how overfitting of

data can be visualized

and here we have a good fit line over

the same data points. So this is also

known as the regression line. Now L2

regularization is also known as ridge

regression. So ridge regression modifies

the overfitted model by adding the

squared magnitude of coefficient as a

penalty term to the loss function. So on

the right you can see a set of data

points plotted and we have our linear

regression line and here we are

calculating the cost function for the

ridge regression line. So our cost

function is actually loss plus lambda

into summation of w ^ 2 where loss is

actually the sum of squared errors or

squared residuals. Lambda stands for

penalty for the errors. W is called the

slope of the curve or line. Okay. Now

consider a case where there are two

points passing through the linear

regression line. Now if you calculate

the cost function, we get the value as

1.69. So here we have assumed that loss

is zero since the two points lie

directly on the line. We have taken

lambda to be 1 and w is 1.3. So if you

use this function or this formula, you

get the cost function as 1.69.

Now moving ahead, let's consider another

situation where we'll calculate the same

cost function for the ridge regression

line. There is some loss for both the

points as they are not on the same line.

So here you can see the sum of squared

residuals is 0.05 which actually is the

sum of 2 squared. I'm assuming this as 2

and 0.1 for this one.

So if you square both and add it the

value is 05 a lambda is again 1 and w we

have assumed to be 6. Now if you find

the cost function the value is 41. Let's

draw the linear regression line and the

ridge regression line with all the

points we find that the ridge regression

line as the best fit since its cost

function is less. Now coming to the

sixth question.

What are the different methods to split

a tree in a decision tree algorithm?

So there are three methods to split a

decision tree. First we have variance.

So reduction in variance is an algorithm

that is used for continuous target

variables. This algorithm uses the

standard formula variance to choose the

best split. So here you can see the

standard formula variance which is

summation of x that is all the

individual points minus xar which is the

mean squared divided by the total number

of observations. Now the split with

lower variance is selected as the

criteria to split the population.

Now the steps to calculate variance is

you need to calculate variance for each

node and then you need to calculate for

each split as the weighted average of

each node variance.

Moving ahead, the second method we have

is information gain. So information gain

is used for splitting the nodes when the

target variable is categorical.

It works on the concept of entropy. Now

the degree of disorganization in a

system is known as entropy. So here you

can see the formula for information gain

which is 1 minus entropy.

Finally we have genie impurity. So genie

impurity is the probability of

incorrectly classifying a randomly

chosen element in the data set if it

were randomly labeled according to the

class distribution in the data set. So

below you can see the formula for gen

impurity. So we have 1 minus summation

of pi whole square where n represents

the number of classes and p of i

represents the probability of randomly

picking an element of class i.

Now moving to the seventh question.

So the question is how do we find the

optimum cluster value in K means

clustering algorithm. Now there are two

methods to find the optimum cluster

value. So first we have the elbow method

which is one of the most wellknown for

finding the optimum number of clusters.

So in this method you need to calculate

the within cluster sum of squared errors

for different values of K and choose the

K for which within cluster sum of

squared errors first starts to diminish.

So in the below plot of squared errors

versus the number of clusters K you can

see at K is equal to 4 the squared error

starts to diminish. So hence our optimum

K value is four. Next we have the siloid

method.

So the celloid method measures how

similar a point is to its own cluster

compared to other clusters. The average

seloid method computes the average of

observations for different values of K.

The optimum number of clusters K is the

one that maximizes the average over a

range of possible values for K. The

seloid score reaches its global maximum

at the optimal K. So in our case the

average seloid reaches maximum at k is

equal to two which you can see here. So

our optimum cluster value will be two

here. Moving ahead the eighth question

in our list is how does the pooling

layer work in a convolutional neural

network. So the pooling layer performs a

downsampling operation in order to

reduce the dimensionality of the feature

map. So in the pooling operation, you

slide a two-dimensional filter over each

channel of feature map and summarize the

features lying within the region covered

by the filter.

It is a common practice to periodically

insert a pooling layer in between

successive convolutional layers in a

convolutional neural network

architecture.

So the pooling layer operates

independently on every depth slice in

the input and resizes it specially using

the max operation. So in the diagram

shown here you can see we have a

rectified feature map. We are using a 2

+2 filter and performing a max pooling

operation. So consider this as the

filter. If you perform the max operation

over the values

let's say 0 5 3 and 1. So considering

this one our pool feature map maximum

value will be five. Similarly for this

chunk of data it is going to be seven.

Next, if you slide the filter over this

square frame, you get eight. And

similarly here you get six. So this is

also known as a pulled feature map.

Moving ahead, the ninth question in our

list is how does LSTM network work? So

long short-term memory networks are a

type of recurrent neural networks that

are capable of learning order dependence

and sequence prediction problems. So

remembering information for long periods

of time is practically their default

behavior. Now, LSTMs also have this

chain-like structure which you can see

here.

But the repeating module has a different

structure. So, instead of having a

single neural network layer, there are

four interacting in a very special way.

Now, you can see these are called as

gates. These gates contain sigmoid

activations. A sigmoid activation is

similar to the tanh activation. Instead

of squishing values between minus1 and +

one, it squishes values between 0 and

one. An LSTDM has four gates.

Now these are called forget, remember,

learn and use or output. So if you see

this in the first step, we use the

forget gate that decides what

information should be thrown away or

kept. the information from the previous

hidden state and the information from

the current input is passed through the

sigmoid function. Values come out

between zero and one.

So if the value is closer to zero, it

means you need to forget that

information and if the value is closer

to one, it means you need to keep that

information.

Next we have the input gate. So the

input gate is used to update the cell

state. First, we pass the previous

hidden state and the current input into

a sigmoid function

that decides which values will be

updated by transforming the values to be

between 0 and 1. Zero means not

important and one means important. You

also pass the hidden state and current

input into the tan function to flatten

the values between minus1 and + one.

This helps to regulate the network.

Then you multiply the tan output with

the sigmoid output. The sigmoid output

will decide which information is

important to keep from the tanh output.

And finally in step three we have the

output gate. This output gate is used to

decide what the next hidden state should

be. First we pass the previous hidden

state and the current input into a

sigmoid function. Then we pass the newly

modified cell state into the tanage

function. We then multiply the tanage

output with the sigmoid output to decide

what information the hidden state should

carry. The output is the hidden state.

The new cell state and the new hidden

state is then carried over to the next

time step. Finally, talking about the

last question in our list of interview

questions, we have explained the concept

of gradient descent in deep learning.

Now gradient descent is an optimization

algorithm which is mainly used to find

the minimum of a function in machine

learning. Gradient descent is used to

update the parameters in a model.

Parameters can vary according to the

algorithms such as coefficients in

linear regression and weights in neural

networks.

You can see we have these maps and on

the y-axis we have the loss. On the

x-axis we have the weight and here we

are trying to find the local minimum or

the global minimum. Now this gradient

descent method is used to minimize the

cost function and update the parameters

of the learning model. The gradient

always points in the direction of the

steepest increase in the loss function.

The gradient descent algorithm takes a

step in the direction of the negative

gradient in order to reduce the loss as

quickly as possible. To determine the

next point along the loss function

curve, the gradient descent algorithm

adds some fraction of the gradient's

magnitude to the starting point. Now

this process is repeated to find the

global minimum.

>> Welcome to math refresher probability

and statistics.

In this lesson, we are going to explain

the concepts of statistics and

probability.

Describe conditional probability. Define

the chain rule of probability. Discuss

the measure of variance. Identify the

types of gshian distribution.

Basic of statistics and probability.

Probability and statistics. Data science

relies heavily on estimates and

predictions. A significant portion of

data science is made up of evaluations

and forecast.

Statistical methods are used to make

estimates for further analysis.

Probability theory is helpful for making

predictions. Statistical methods are

highly dependent on probability theory

and all probability and statistics are

dependent on data.

Data is information acquired for

reference or research via observations,

facts, and measurements. Data is a set

of facts structured in the form that

computers can interpret such as numbers,

words, estimations, and views.

Importance of data. Data aids in seeing

more about the information by

identifying possible connections between

two features. Data assists in the

detection of distortion by uncovering

hidden patterns based on prior

information patterns. Data may be

utilized to anticipate the future or

predict the current state of affairs.

Also, data aids in determining whether

two pieces of information have any

instance in common or not. Types of

data. Data might be quantitative. That

is data that can be measured or counted

in numbers or it may be qualitative

which is data which is generally divided

into groups or in simpler words which

cannot be counted or measured in

numbers. Let's consider an example a

customer information data of a bank may

contain quantitative and qualitative

data. Consider this snapshot where we

have customer ID, surname, geography,

gender, age, balance, has C or card is

active member. Amongst these variables

we can see surname is mostly qualitative

as it cannot be counted and measured in

numbers. Geography and gender are also

qualitative as they cannot be counted in

numbers and are mostly groups. has C or

card that is has credit card and is

active member although are containing

numerical in form but these are

categorical that means these have been

divided into groups of one and zero that

represent yes and no as an answer hence

these two variables are also qualitative

customer ID is again although a

numerical data however the significance

or intuition behind Customer ID is

categorical.

Hence, it may be kept in the qualitative

data also. However, age and balance

these are numerical information which

have been measured or counted and

numerical operations can be performed on

them. Hence, these are under

quantitative data categories.

Introduction to descriptive statistics.

Descriptive statistics. A descriptive

measurement is summary measure that

quantitatively portrays the most

important features of a set of data

allowing for a better comprehension of

the information. Data can be measured as

different levels. The levels of

measurement describe the nature of

information stored in the data assigned

to the variables. Qualitative data can

be measured as nominal or ordinal.

Quantitative data can be measured in

terms of interval and ratio type.

Nominal data. The data is categorized

using names, labels or qualities. For

example, brand name, zip code, and

gender. Ordinal data can be arranged in

order or ranked and can be compared.

Examples include grades, star reviews,

position, and race, and date. Interval

data is the data that is ordered and has

meaningful differences between the data

points. Example temperature in Celsius

and year of birth. Ratio data is similar

to the interval level with the added

property of inherent zero. Mathematical

calculations can be performed on both

interval as well as ratio data. For

example, height, age, and weight.

Population versus sample. Before

analyzing the data, it's important to

figure out if it's from a population or

a sample. Population is a collection of

all available items as well as each unit

in our study. Sample is a subset of the

population that contains only a few

units of the population. Population data

is used for study when the data pool is

very small and can give all the required

information. Samples are collected

randomly and represent the entire

population in the best possible way.

Measures of central tendency.

The central tendency is a single value

that aids in the description of the data

by determining its center position.

Measures of central tendency are

sometimes known as summary statistics or

measures of central location.

The most popular measurements of central

tendency are mean, median, and mode. The

normal distribution is a bell-shaped

symmetrical distribution in which mean,

median, and mode all are equal. The

curve over here shows the bell-shaped

curve or the normal distribution of

variable X. The point over here that is

X1 is the point which represents the

mean, median and mode of this

distribution. Mean mean is calculated by

dividing these sum of all data values by

the total number of data values. It gets

affected when there are unusual or

extreme values. It is sensitive to the

outliers. Mean can be calculated as

summation over all the values of X in a

collection divided by the size of the

collection.

For example, we have a collection where

we have values as 7 3 4 1 6 and 7.

We find out the sum of these values

which is 28 and there are total of six

values. So 28 / 6 gives us a mean value

of 4.66.

Median,

it is the middle value in the set of the

data that has been sorted in ascending

order.

It is a better alternative to mean since

it is less impacted by outliers and

skewess.

It is closer to the actual central

value.

Median is calculated differently for

different sizes of data.

Differentiated as if the total number of

values is odd or if the total number of

values is even. If the size of the data

is odd. For example, in this case we

have five elements.

After sorting whatever middle value we

get

that means n + 1 by 2 term in this case

5 + 1 / 2

that is the third term which is four is

the median value.

In case when the total number of values

is even like here there are six values.

The average or the mean of the two

central values is considered as the

median. In this case the median is the

mean of 6 and four which is five. Mode.

Mode represents the most common value in

the data set. It is not at all affected

by extreme observations.

It is the best measure of central

tendency for highly skewed or non-normal

distribution.

Mode for categorical data is determined

by estimating the frequencies for each

categories

and then the category with the highest

frequency is considered to be mode.

Like in this case 7 has the highest

frequency. Hence seven becomes the mode

value. However, in case of continuous

data or quantitative data, the

calculation of mode is slightly

different. The first step in calculation

of mode is dividing the data into

classes which are equal with then

getting the frequency of data points

lying in within that range of classes

and finally selecting the class with the

highest frequency.

Using the range of that class and the

frequencies, we can get the final mode

value.

Using the formula L+

minus F_sub_1 multiplied to H / FM minus

F_sub_1 plus FM minus F_sub_2.

Here L is the lower limit or the lower

observation of the mode class.

H is the size of the mode class.

FM is the frequency of the mode class.

F_sub_1 is the frequency of the class

proceeding to mode and F_sub_2 is the

frequency of the class succeeding to

mode. This gives us the final mode

value,

mean versus expectation.

Now let's talk about mean versus

expectation.

So in general we use the expected value

or expectation when we want to calculate

the mean of a probability distribution

that represents the average value we

expect to occur before collecting any

data. And mean on the other hand mean is

basically used when we want to calculate

the average value of a given sample.

This represents the average value of raw

data that we may have already collected.

We can understand this by using a simple

example.

Now to calculate the expected value of

this probability distribution, we can

use a specific formula from the previous

discussion.

This is going to be the expected value

where X is going to be the data value

and this PX is the probability of value.

For example, we could calculate the

expected value for this probability

distribution to be as shown.

So here it will be 1.45 goals.

So this represents the expected number

of goals that the team will score in any

given game.

And then if you talk about calculating

mean, so we typically calculate the mean

after we have actually collected raw

data.

For example, suppose we record the

number of goals that a soccer team will

score in 15 different games.

Now to calculate the mean number of

goals scored per game,

we can use the following formula

where sum of x is basically the sum of

all the goals divided by n and the

number of records or we can say the

sample size.

It is as shown on the screen.

So this represents the mean number of

goals scored per game by the team.

Measures of asymmetry.

The difference between the three

distinct curves can be studied in this

image.

The central curve is the normal or no

skewess curve. Here mean, median and

mode all lie on the same point. This

normal curve is symmetrical about its

mean, median and mode.

That means the left hand side of the

curve is a mirror image of the right

hand side of the curve.

However, in case of negatively skewed

data, the tail is elongated on the left

hand side

and the mean is smaller than the mode

and the median values or is on the left

hand side of the mode.

Hence indicating that the outliers are

in the negative direction.

On the other hand, in case of positively

skewed, the data is concentrated on the

left hand side of the curve.

While the tail is elongated or longer on

the right hand side of the curve,

the mean is greater than the mode and

median

or is on the right hand side of the mode

and median indicating that the outliers

are in the positive direction.

Let's consider an example.

The graph here shows the global income

distribution for the year 2003 2013 and

a projection for 2035.

If we see the global income distribution

statistics for 2003 it is highly right

skewed.

We can observe in the previous graph

that in 2003

the mean of $3,451

was higher than the median of $1090.

The global income is definitely not

evenly distributed. The majority of

people make less than $2,000 each year.

while only a small percentage of the

population earns more than $14,000.

Measures of variability.

Measures of variability.

Dispersion. The measure of central

tendencies provide a single value that

addresses the full worth. However, the

central tendency cannot depict the

viewpoint entirely. The metric of

dispersion helps us focus on the

inconsistency in the data spread.

Measures of dispersion describe the

spread of the data.

The range, intercortile range, standard

deviation and variance are examples of

dispersion measures.

Range.

The range of distribution is the

difference between the largest and the

smallest amount of data.

The range, for example, does not include

all of a series positive aspects.

It concentrates on the most shocking

aspects and ignores that aren't

considered critical. For example, for a

set 13, 33, 45, 67, 70.

The range is 57. That is the maximum of

this which is 70 minus the minimum over

here which is 13.

Variance.

Variance is the average of all squared

deviations.

It is defined as the sum of squared

distance between each point and the mean

or the dispersion around the mean.

The standard deviation is used as

variance suffers from a unit difference.

Variance can be computed as sigma square

summation over x - mu^ 2

divided by n

where mu is the mean of the data, x is

the individual data point

and n is the size of the data.

This representation is for a population

data.

For a sample data variance can be

computed as X minus

Xar whole square summation

over it divided by n minus one.

Here Xbar is the mean of these sample

data and n is the sample size.

The units of values and variance are not

equal.

So another variability measure is used.

Standard deviation.

Standard deviation is a statistical term

used to measure the amount of

variability or dispersion around a mean.

The standard deviation is calculated as

the square root of variance. It depicts

the concentration of the data around the

mean of the data set.

Standard deviation as indicated

previously can be computed as square

root of variance

for a population data. Standard

deviation sigma can be computed as

square root of summation over x i minus

mu^ square / n

where mu is the mean of the data x i are

the data points and n is the size. Let's

consider an example.

Let's find out the mean, variance, and

standard deviation for this data. The

data values are three, 5, 6, 9, and 10.

To find out the mean, we first find the

sum of all these data values

that is 33 and divide it by the count,

which is five.

We get the mean of 6.6. To compute the

variance, we start by computing the

deviation.

That is X minus the mean of X. Here

three is one of the values of the data

and 6.6 is the mean.

So 3 - 6.6 squared and we do that

to find out sum of all the deviations

divided by the count

which is five.

we end up getting an overall variance of

6.64.

Standard deviation as we know is

measured at square root of variance that

is square of 6.64

which amounts to 2.576.

Measures of relationship.

Measures of relationship. Coariance.

Coariance is the measure of joint

variability of two variables.

It measures the direction of the

relationship between the variables. It

determines if one variable will cause

the other to alter in the same way.

Coariance between variable X and Y can

be computed as summation over the

product of X I - XR

and Y I - Y bar the whole divided by N

minus one.

Here Xar and Y bar are the mean of X and

Y respectively. The value of covariance

can range from minus infinity to a plus

infinity.

Correlation.

Correlation is normalized coariance.

It measures the strength of association

between two variables. The most common

measure for correlation is the Pearson

correlation coefficient.

Correlation between two variables

X and Y can be measured with respect to

coariance as coariance between X

and Y divided by the standard deviation

of X and standard deviation of Y.

The value of correlation ranges from a

negative 1 to positive 1.

Types of correlation.

Correlation can be either a positive

correlation,

zero correlation or a negative

correlation.

The first picture over here represents a

perfect positive correlation

wherein a straight line with a positive

slope

is representing the relationship between

the two variables.

Zero correlation means that the line

representing the relationship between

the two variables is horizontal to the

xaxis.

Perfect negative correlation can be

represented by a straight line with a

negative slope.

Correlation equals to 1 implies a

positive relationship. That is when one

variable increases the other variable

also increases. A correlation value of

negative 1 implies a negative

relationship. That is when one variable

increases the other decreases.

The correlation coefficient of zero

shows that the variables are completely

independent of each other.

Let's consider an example.

Here we have two variables height and

weight.

To compute the correlation between

height and weight,

we use the correlation formula as

covariance of X

and Y divided by standard deviation of X

and standard deviation of Y.

Here height is the X variable and weight

is the Y variable.

First to compute coariance we compute

the x - xar and y - y bar values and

then the product of them.

We then compute x - xrยฒ

and y - y bar square values to compute

the standard deviations of height and

weight respectively. Correlation as we

know has been defined as covariance of x

and i and y divided by standard

deviations of x and y.

This can also be represented as

summation over x - xr multiplied to y -

y bar

divided by square root of summation over

sum of squared deviations

that is x - xr square multiplied to

square root of summation over y - yar

whole square that is sum of square

deviations for y.

Now let's find out values to put into

this formula.

First we find out the overall sum of

height to get the mean of height which

is 5.14.

Similarly we get the sum of weight to

get the mean of weight as 50. We now get

the summation over x - xr multiplied to

y - y bar to get the numerator for the

formula. Then we compute x - xr square

summation

and y - y bar square that is sum of

squared deviation of x and y

respectively.

Now we put in the values in this final

correlation formula to get a correlation

value of 0.889.

This indicates that height and weight

have a positive relationship.

It is evident that as height grows,

weight also increases.

In this module, we will be talking about

expectation and variance.

So the expected value or we can say mean

of a given variable that we can denote

by X is a discrete random variable where

it is a weighted average of the possible

values that X can take and each value is

going to be according to the probability

of that specific event occurring.

So usually the expected value of X is

denoted by a simple formula where we can

define the expectation based on the X

parameter.

which is going to be the sum of each

possible outcome multiplied by the

probability of the outcome occurring.

So in more concrete terms, the

expectation is what we would expect the

outcome of an experiment to be on

average.

We can take an example for the coin. If

a coin is being tossed 10 times, then

one is most likely to get five heads and

five tails.

Same logic can be discussed if we talk

about another example of rolling a die.

So there are six possible outcomes when

you roll a dieice 1 2 3 4 5 6. And each

of these has a probability of 1 by 6 of

occurring. So we can say that the

expectation is going to be 1 multiplied

by the probability of that happening

which is going to be 1x 6 + 2x 6 + 3x 6

+ 4x 6 + 5x 6 + 6x 6 and that is going

to give us 3.5 as an output. The

expected value is 3.5.

So if you think about it, 3.5 is halfway

between the possible values that I can

take and this is what we should have

expected.

Next we talk about the concept of

variance. So variance of a random

variable allows us to know something

about the spread of the possible values

of the variable. So for a discrete

random variable X the variances of X is

going to be denoted by using a simple

formula that is going to be var=

E X - M the whole square where M is

basically the expected value of the

expectation of X. So this is more like a

standard deviation of X which can also

be represented by using this formula. So

the variance does not behave in the same

way as expectation when we multiply and

add constants to random variables.

So now there are two different type of

variance that we can have a fair

understanding on. First of all we have

low variance and then we have high

variance.

So low variance simply means that there

is a small variation in the production

of the target function with changes in

the trading data set and at the same

time high variance as we can see here

high variance shows a large variation in

prediction of the target function with

changes in the trading data set. So a

model that shows high variance learns a

lot and perform well with the training

data set and it does not generalize well

with the unseen data set and that's why

as a result such a model gives good

results with training data set but shows

high error rates on the test data set

and since the high variance a model

learns too much from the data set it

leads to an overfitting of the model. So

model with high variance will be having

couple of issues like it may lead to

overfitting or it may also lead to

increase in model complexities.

Next we have skewess.

So skewess in simple terms is basically

a measure of asymmetry of a

distribution. So distribution is

asymmetrical when its left and right

sides are not the mirror images.

Right now this is a mirrored image and a

distribution can have right positive or

we can say negative or it can have zero

skewess.

So right skewed in this scenario is

basically the distribution is longer on

the right side of its peak

and a left skew distribution is going to

be we can say where it is longer on the

left side.

So we can see we have this one as a part

of right side. It is more elongated

towards the right side and this one is

more elongated towards the left side. So

we can think of skewess in terms of

tails. A tail is long tampering and the

end of a distribution. So it simply

indicates that they are observations at

one end of the distribution but that

they are relatively infrequent. So a

right skew distribution has a long tail

on the right side as you can see here.

So the number supports observed. Let's

say we have a data on a per year basis.

So again we can have a more skewess

towards the right side where data is

being dropping as we continue to

increase the number of years. For

example we may have a high sales towards

the beginning of year suppose in 2022

but again as we proceed to 2023 second

half we are seeing the dip in

performance. So that is rightly skewed

and same way let's suppose if we started

with the sales figure it was really less

in suppose 2002

but again as we proceeded to 2023 now

our sales have been gradually

increasing. So it's more like skew

towards the left section as a part of

negative skew. Next we have curtosis.

So curtosis is basically a measure of

the tailness of a distribution.

So taeness is how often the outliers

occur and act as curtis is the tailness

of the distribution related to a normal

distribution. So a distribution with

medium curttosis is called as messortic.

A distribution with low curtosis like

this one. This is called as the

platicurtic and then distribution with

high curtosis like this one. This is

called as the leptocortic.

So tails here they are tapering ends on

either side of a distribution like this.

So they represent the probability or the

frequency of values that are extremely

high or extremely low to the mean.

In other words, tails here represents

how often the outliers occur.

So there are three type of curtis. We

have platicurtic which is negative,

leptocortic which is a positive towards

the upper end and then we have messertic

which is a normal distribution. So

messertic is the medium tail. So normal

distributions they have a curtosis of

three. So any distribution with a

kurtosis of approx value of three is

going to be messertic. And curtosis is

described in terms of excess curttosis

which is curtosis minus3. And since

normal distribution they have a curtosis

of three axis curtises makes comparing a

distribution curtosis to a normal

distribution even easier. Introduction

to probability.

Probability theory. Probability is a

measure of the likelihood that an event

will occur.

Let's consider an example of coin toss

where the chances of getting heads on a

coin are 1 by two or 50%.

The probability of each given event is

between zero and one both inclusive. Sum

of an events cumulative probability

cannot be greater than one.

Hence the probability of an event X lies

between zero and one. This means that

the integral of probability of

distribution over x equals to 1.

Conditional probability. Conditional

probability of any event A is defined as

the probability of occurrence of A given

that event B has previously occurred.

Condition probability of event A given B

can be estimated as probability of A

intersection B that is probability of

both A and B happening together

divided by the probability of B.

It is also written as that probability

of A intersection B equals to

probability of A given B multiplied to

probability of B.

Let's consider an example.

In a coin, we are doing a two coin flip.

Coin one gets heads, tails, heads, and

tails in subsequent flips.

while coin two gets tails, heads, heads,

and tails in the subsequent flips. Now,

the probability that coin one will get a

head is 2 out of four. While the

probability that coin two will get heads

is again two out of four.

The probability that both coin one and

coin two will have a heads is just one

out of the four flips.

Hence the probability that coin one will

get heads given that coin 2 is already

heads can be computed as probability of

coin one edge intersection coin 2 edge

that is 1x4 divided by probability of

coin 2 edge

that's a given that is 2x 4 which is

going to be 0.5 or 50% based

base theorem Base theorem calculates the

conditional probability of an event

based on its prior probabilities.

Basically base theorem incorporates the

prior probability distribution to

predict the posterior probabilities.

Base theorem for conditional probability

can be expressed as probability of A

given B equals probability of B given A

divided by probability of B multiplied

to probability of A.

Base theorem allows updating the

probability values by using new

information or evidence. Here

probability of A is known as prior

probability. That is the probability of

event before any new data is collected.

Probability of A given B is known as the

posterior probability. It is the revised

probability of an event occurring after

taking into consideration the new

information probability of B given A is

known as the likelihood and probability

of B is probability of observing an

evidence B model. An example consider an

example for calculating the likelihood

of having diabetes based on frequency of

fast food consumption. Here is the

observed data. Let's say the fast food

audience is 20%. Diabetes prevalence is

10% and 5% is fast food and diabetes.

The chances of diabetes given fast food

that is the conditional probability of D

given B can be calculated as probability

of diabetes and fast food together

divided by probability of fast food.

That means 5% divided by 20%. that

equals 25%.

Define an analysis can state eating fast

food increases the chance of having

diabetes by 25%.

The multiplication rule of probability

if events A and B are statistically

independent and probability of A

intersection B can be given as

probability of A given B multiplied to

probability of B. However, probability

of A intersection B is also given as

probability of A multiplied to

probability of B. Here probability of A

given B equals to probability of A when

we assume that probability of B is non

zero. Similarly, probability of B equals

probability of B given A assuming

probability of A is non zero.

Chain rule of probability joint

probability distributions over many

random variables can be reduced into

conditional distributions over a single

variable. It can be expressed as

probability of X1 X2 so on until Xn

equals probability of X1 intersection

probability of X I given probability of

X1 till X I minus one.

For example, the joint probability of A,

B and C can be given as probability of A

given B. C multiplied to probability of

B given C multiply to probability of C.

Logistic sigmoid.

The logistics function is a type of

sigmoid function that aims to predict

the class to which a particular sample

belongs. Its outcome is discrete binary

value. a probability between zero and

one. The logistic sigmoid is a useful

function that follows the yes curve. It

saturates when the input is very large

or very small. Logistic sigmoid is

expressed as sigma of x= 1 upon 1 + e to

the power minus x.

The logistic sigmoid can be expressed as

sigmoid function of x is given as 1 upon

1 + e ^ minus x where e is the ooler's

number.

Gshian distribution.

The gossian distribution is a type of

distribution in which data tends to

cluster around a central value with

little or no bias to the left or right.

It is often referred to as normal

distribution.

In absence of prior information, the

normal distribution is frequently a fair

assumption in machine learning

equation.

The formula for calculating Gaussian

distribution is described as the normal

distribution of X.

That is the function of x given mean as

mu and variance is sigma square can be

calculated as 1 upon sigma square

roo of 2 pi e to the power -/ x -

mood / sigma square

where mu is the mean or peak value which

also is the expected value of x.

Sigma is the standard deviation. Sigma

square is the variance.

A standard normal distribution has a

mean of zero and a standard deviation of

one.

Gshian distribution can be univariate

which describes the distribution of a

single variable X.

It can also be multivariate where it can

just use to describe the distribution of

several variables.

It is represented in 3D of ND formats.

Law of large numbers.

Now let's talk about law of large

numbers. The law of large numbers states

that an observed sample average from a

large sample will be close to the true

population average and that it will get

closer in the larger sample. So the law

of large number does not guarantee that

a given sample spatially a small sample

will reflect the true population

characteristics or that a sample does

not reflect the true population will be

balanced by a subsequent sample. This is

for the law of large numbers to express

the relationship between scale and

growth rate.

So there are multiple examples through

which we can understand

and it is widely used in statistical

analysis in working with the central

limit theorem in terms of the business

growth. So there are multiple real time

setup in which these are going to be

used. So if you talk about tossing a

coin so tossing a coin in a number of

times will give us two different type of

outcomes.

the result will spread evenly between

head and tails and the expected average

value is going to be half.

That means 50 times tails and 30 times

heads. But again, if you toss a coin

1,000 times, then the result can be in

different manners because out of 1,000,

let's say 850 times it has been head and

only 150 times it has been tails and so

on. So that's why the possibility of one

event occurring is going to be changed

in large sample sets as compared to a

small sample sets as in let's say 10

times. So the number of heads and tails

unbalanced for lower number of trials.

So we can see it is unbalanced.

But again as soon as we toss more number

of coins more leans towards the balance

value or we can see the observed

averages.

Next we have p value.

So p value is basically a number

calculated from the statistical test

that describes how likely we are to have

found a particular set of observations

if the null hypothesis were true. So p

values are used in hypothesis testing to

help decide whether to reject the null

hypothesis. And the smaller the p value,

the more likely we are to reject the

null hypothesis.

So we have a term called as null

hypothesis. So all statistical tests

they have null hypothesis. So for most

tests the null hypothesis is that there

is no relationship between our variables

of in first or that there is no

difference among groups. For example in

a two-tail t test the non-hypothesis is

that the difference between two groups

is going to be zero.

So p value is going to tell us how

likely it is that our data could have

occurred under the null hypothesis.

It is done by calculating the likelihood

of a test statistic

which is the number calculated by a

statistical test using our data. So p

value tell us how often we would expect

to see a test statistic as extreme or

more extreme

than one calculated by a statistical

test. if the null hypothesis of the test

was true.

So there are multiple limitations as

well. So first one is the results can be

significant but again they are they may

not be practical as we have compared it

can be based on multiple hypothesis for

a game for the healthcare test. If the

test is going to be positive or not it

may show even values of the effect of a

variable but not the magnitude in real

life. What exactly is going to be the

application of a drug test being failed

in pharma company? Therefore, it is

recommended to use confidence and levels

in addition to the p values to quantify

or we can say to give a solid figure to

the reserve which we are going to get.

The p values they are interpreted as

supporting or we can say refuting the

alternative hypothesis.

So p value can only tell you whether or

not the null hypothesis is supported. It

cannot tell us whether our alternative

hypothesis is true or why. So the risk

of rejecting the null hypothesis is

often higher than the p value. So

especially when we are looking at a

single study or when using small sample

sizes. So this is because the smaller

frame of reference, the greater are the

chance that as we stumble across a

statistically significant pattern

completely by accident.

Key takeaways.

Key takeaways. Probability and

statistics structure the premise of the

data. The data helps in anticipating the

future or gauging in view of the past

patterns of information.

The central tendency is a single value

that helps to describe the data by

identifying these central positions. The

mean, median, and mode are the measures

of central tendencies.

The distribution where the data tends to

be around a central value with a lack of

bias or minimal bias towards the left or

right is called as gshian distribution.

>> My name is Richard Kersner with the

simply learn team. That's get certified,

get ahead. We're going to cover

mathematics for machine learning. So

today's agenda is going to cover data

and its types. Then we're going to dive

into linear algebra and its concepts,

calculus, statistics for machine

learning, probability for machine

learning, hands-on demos, and of course

throwing in there in the middle is going

to be your matrixes and a few other

things to go along with all this.

Data and its types. Data denotes the

individual pieces of factual information

collected from various sources. It is

stored, processed and later used for

analysis.

And so we see here uh just a huge

grouping of information, a lot of tech

stuff, money, dollar signs, numbers

uh and then you have your performing

analytics to drive insights and

hopefully you have a nice share your

shareholders gathered at the meeting and

you're able to explain it in something

they can understand. So we talk about

datas types of data we have in our types

of data we have a qualitative

categorical

you think nominal or ordinal and then

you have your quantitative or numerical

which is discrete or continuous

and let's look a little closer at those

data type vocabulary always people's

favorite is the vocabulary words okay

not mine uh but let's dive into this

what we mean by nominal nominal they are

used to label various just uh label our

variables without providing any

measurable value. Uh country, gender,

race, hair, color, etc. It's something

that you either mark true or false. This

is a label. It's on or off. Either they

have a red hat on or they do not. Uh so

a lot of times when you're thinking

nominal data labels, uh think of it as a

true false kind of setup. And we look at

ordinal. This is categorical data with a

set order or a scale to it. Uh and you

can think of salary range is a great

one. Uh movie ratings etc. You see here

the salary range if you have 10,000 to

20,000 number of employees earning that

rate is 150 20,000 to 30,000 100 and so

forth. Some of the terms you'll hear is

bucket. Uh this is where you have 10

different buckets and you want to

separate it into something that makes

sense into those 10 buckets. And so when

we start talking about ordinal, a lot of

times when you get down to the brass

bones, again, we're talking true false.

Uh so if you're a member of the 10 to

20k range, uh so forth, those would each

be either part of that group or you're

not. But now we're talking about buckets

and we want to count how many people are

in that bucket. Quantitative numerical

data uh falls into two classes, discrete

or continuous. And so data with a final

set of values which can be categorized

class strength questions answered

correctly and runs hit in cricket. A lot

of times when you see this you can think

integer uh and a very restricted integer

i.e. you can only have 100 questions um

on a test. So you can it's very

discreet. I only have a 100 different

values that it can attain. So think

usually you're talking about integers

but within a very small range. They

don't have an open end or anything like

that.

Uh so discrete is very solid, simple to

count, set number. Continuous on the

other hand uh continuous data can take

any numerical value within a range. So

water pressure, weight of a person etc.

Usually we start thinking about float

values where they can get phenomenally

small in their in what they're worth.

And there's a whole series of values

that falls right between discrete and

continuous. Um you can think of the

stock market. You have dollar amounts.

It's still discreet, but it starts to

get complicated enough when you have

like, you know, jump in the stock market

from $525.33

to $580.67.

There's a lot of point values in there.

It'd still be called discreet, but you

start looking at it as almost continuous

because it does have such a variance in

it. Now uh we talk about n we did we

went over nominal and ordinal uh almost

true false charts and we looked at

quantitative and numerical data which

we're starting to get into numbers.

Discrete you can usually a lot of times

discreet will be put into it could be

put into true false but usually it's

not. Uh so we want to address this stuff

and the first thing we want to look at

is the very basic which is your algebra.

So we're going to take a look at linear

algebra. You can remember back when your

uklidian geometry uh we have a line.

Well, let's go through this. We have

linear algebra is the domain of

mathematics concerning linear equations

and their representations in vector

spaces and through matrices. I told you

we're going to talk about matrices. Uh

so a linear equation is simply um uh 2x

+ 4 y - 3 z = 10. Very linear. 10 x +

12.4 4 y = z. And now you can actually

solve these two equations by combining

them. Uh, and that's we're talking about

a linear equation.

In the vectors, we have a + b= c. Now,

we're starting to look at a direction.

And these values usually think of an xyz

plot. Um, so each one is a direction.

And the actual distance of like a

triangle A is C. And then your matrix

can describe all kinds of things. Um, I

find matrixes uh confuse a lot of

people, not because they're particularly

difficult, but because of the magnitude

and the different things are used for.

And a matrix is a chart or a um, you

know, think of a spreadsheet, but you

have your rows and your columns. And

you'll see here we have a * b= c. Very

important to know your counts. Uh, so

depending on how the math is being done,

what you're using it for, making sure

you have the same rows and the number of

columns or a single number, there's all

kinds of things that play in that that

can make matrixes confusing. Uh, but

really it has a lot more to do with what

domain you're working in. Uh, are you

adding in multiple polomials where you

have like uh uh ax^2 plus b y plus, you

know, you start to see that can be very

confusing versus a very straightforward

matrix. And let's just go a little

deeper into these because these are such

primary this is what we're here to talk

about is these different math uh

mathematical computations that come up.

So we're looking at linear equations.

Let's dig deeper into that one. An

equation having a maximum order of one

is called a linear equation. Uh so it's

linear because when you look at this we

have uh ax plus b= c which is a one

variable. We have two variable ax plus b

y = c ax plus b y plus z c cz z= d and

so forth. But all of these are to the

power of one. You don't see x squar. You

don't see x cubed. So we're talking

about linear equations. That's what

we're talking about in their addition.

If you have already dived into say

neural networks, you should recognize

this ax plus b y plus cz um setup plus

the intercept. uh which is basically

your your neural network each node

adding up all the different inputs and

we can drill down into that most common

formula is your y = mx + c.

So you have your uh y equals the m which

is your slope, your x value plus c which

is your um y intercept. They kind of

labeled it wrong here

threw me for a loop but the the c would

be your y intercept. So when you set x

equal to zero, y equals c. And that's

that's your y intercept right there. Uh

and that's they they just had reversed

value of y. When x equals 0, it equals

the y intercept, which is c. And your

slope gradient line, which is your m. So

you get your y = 2x + 3. And there's

lots of easy ways to compute this. This

why this is why we always start with the

most basic one when we're solving one of

these problems. And then of course the

um one of the most important takeaways

is the slope gradient of the line. Uh so

the slope is very important that m

value. Uh in this case we went ahead and

solved this. If you have y = 2x + 3 you

can see how it has a nice line graph

here on the right.

So matrixes a matrix refers to a

rectangular representation of an array

of numbers arranged in columns and rows.

So we're talking m rows by n columns

here. A11 is denotes the element of the

first row in the first column. Similarly

a12 and it's really pronounced a11 in

this particular setup. So it's row one

column one. A12 is a of row one column 2

uh first row and second column and so

on.

And there's a lot of ways to denote

this. I've seen these as like a capital

letter A, smaller case A for the top row

or I mean you can see where they can go

all kinds of different directions as far

as the value. You just take a moment to

realize there's need to be some

designation as far as what row it's in

and what column it's in. And we have our

uh basic operations. We have addition.

So when you think about addition, you

have uh two matrices of 2x two and you

just add each individual number in that

matrix and then when you get to the

bottom you have uh in this case the

solution is 12, 10 + 2 is 12, 5 + 3 is 8

and so on. And the same thing with

subtraction.

Now again you're counting matrices you

want to check your um dimensions of the

matrix the shape you'll see shape come

up a lot in programming. So we're

talking about dimensions we're talking

about the shape. If the two shapes are

equal this is what happens when you add

them together or subtract them. And we

have multiplication. When you look at

the multiplication you end up with a

very slightly different setup going.

Now, if we look at our last one, we're

um uh we're like, why? This always gets

to me when we get to matrices. They

don't really say why you multiply

matrices. Um you know, my first thought

is 1 * 2, 4 * 3. But if you look at

this, we get 1 * 2 + 4 * 3, 1 * 3 + 4 *

5,

uh 6 * 2 + 3 * 3, 6 * 3 + 3 * 5. If

you're looking at these matrices, uh,

think of this more as an equation. And

so we have, uh, if you remember when we

back up here for our multiple line

equations, let's just go back up a

couple slides where we were looking at,

uh, two variable. So this is a two

variable equation. ax plus b y= c.

Um, and this is a way to make it very

quick to solve these variables. And

that's why you have the matrix, and

that's why you do

the multiplication the way they do. And

this is the dotproduct of uh 1 * 2 + 4 *

3

1 * 3 + 4 * 5

uh 6 * 2 + 3 * 3 6 * 3 + 3 * 5. And it

gives us a nice little 14, 23, 21, and

33 over here, which then can be used and

reduced down to a simple um formula as

far as solving the variables as you have

enough inputs. Uh and then in matrix

operations, when you're dealing with a

lot of matrices, uh now keep in mind

multiplying matrices is different than

finding the product of two matrices.

Okay? So we're talking about

multiplication, we're talking about

solving uh for equations. When you're

finding the product, you are just

finding one time two. Keep that in mind

because that does come up. I've had that

come up a number of times where I am

altering data and I get confused as to

what I'm doing with it. Uh transpose

flipping the matrix over it's diagonal.

Comes up all the time where you have you

still have 12, but instead of it being

uh 128, it's now 1214 821. You're just

flipping the columns and the rows. Uh

and then of course you can do an inverse

um changing the signs of the values

across this main diagonal. And you can

see here we have the inverse a to the

minus1 and ends up with uh instead of 12

8 14 12 it's now -22 -12 vectors uh

vector just means we have

a value and a direction and we have down

four numbers here on our vector.

uh in mathematics a one-dimensional

matrix is called a vector. Uh so if you

have your xplot and you have a single

value that values along the x- axis and

it's a single dimension. If you have two

dimensions you can think about putting

them on a graph. You might have x and

you might have y and each value denotes

a direction. And then of course the

actual distance is going to be the

hypothesis of that triangle. Uh and you

can do that with three dimensionals x y

and z. uh and you can do it all the way

to nth dimensions. So when they talk

about the k means uh for categorizing

and how close data is together, they

will compute that based on the

Pythagorean theorem. So you would take

uh the square of each value, add them

all together and find the square root

and that gives you a distance as far as

where that point is, where that vector

exists or an actual point value. And

then you can compare that point value to

another one and it makes a very easy

comparison versus comparing uh 50 or 60

different numbers. And that brings us up

to gene vectors and I gene values. Uh I

gene vectors the vectors that don't

change their span while transformation

and I gene values the scalar values that

are associated to the vectors.

Conceptually you can think of the vector

as your picture. you have a picture.

It's um uh two dimensions x and y. And

so when you do those two dimensions and

those two values or whatever that value

is um that is that point but the values

change when you skew it and so if we

take and we have a vector a and that's a

set value uh b is um your is your you

have a and b which is your hygiene

vector. Two is the i gene value. So,

we're altering all the values by two.

That means we're u maybe we're

stretching it out one direction, making

it tall if you're doing picture editing.

Um that that's one of the places this

comes in. But you can see when you're

transforming uh your different

information, how you transform it is

then your hygiene value. And you can see

here uh vector after line transition

uh we have 3 a is the hygiene vector.

Three is the hygiene value. So A doesn't

change. That's whatever we started with.

That's your original picture. And three

uh is skewing it one direction and maybe

uh B is being skewed another direction.

And so you have a nice tilted picture

because you've altered it by those by

the hygiene values.

So let's go ahead and pull up a demo on

linear algebra. And to do this, I'm

going to go through my trusted Anaconda

into my Jupiter notebook. and we'll

create a new uh notebook called linear

algebra. Since we are working in Python,

uh we're going to use our numpy. I

always import that as np or numpy array.

Probably the most popular um module for

doing matrixes and things in

given that this is part of a series. I'm

not going to go too much into numpy. Uh

we are going to go ahead and create two

different variables. A for a numpy array

10 15 and b 29.

We'll go ahead and run this. And you can

see there's our two arrays 105 29. And I

went ahead and added a space there in

between so it's easier to read. And

since it's the last line, we don't have

to put the print statement on it unless

you want. We can simp but we can simply

do a plus b. So when I run this, uh, we

have 10 15 29 and we get 30 24, which is

what you expect. 10 + 20 15 + 9. You

could almost look at this addition as

being um

just adding up the columns on here

coming down. And if we wanted to do it a

different way, we could also do a t plus

b dot t. Remember that t flips them. And

so if we do that, we now get them uh we

now have 304 going the other way. We

could also do something kind of fun.

There's a lot of different ways to do

this. Uh, as far as a plus b, I can also

do a plus b. T and you're going to see

that that will come out the same. The 30

24 whether I transpose a and b or

transpose them both at the end.

And likewise, we can very easily

subtract two vectors. I can go a minus

b. And we run that and we get - 106. Now

remember, this is the last line in this

particular section. That's why I don't

have to put the print around it. Um, and

just like we did before, we can

transpose either the individual or we

can transpose the main setup and then we

get a minus 106 going the other way.

Now, we didn't mention this in our

notes, but you can also do a scalar

multiplication.

Let me just put down scaler so you can

remember that. Uh what we're talking

about here is I have uh this array here

u and if I go a time u uh we'll take the

value two we'll multiply it by every

value in here. So 2 * 30 is 60 2 * 15

and just like we did before

um this happens a lot because when

you're doing matrices you do need to

flip them you get 6030 coming this way.

So in numpy uh we have what they call

dotproduct

and uh what this this in a

twodimensional vectors it is the

equivalent of two matrix multiplication

and remember we were talking about

matrix multiplication

uh where it is the well let's walk

through it

we'll go ahead and start by defining two

um numpy arrays we'll have uh 10 20 256

or our u and our E uh and then we're

going to go ahead and do if we take

the values uh and if you remember

correctly

an array like this would be 10 * 25 + 20

* 6. We'll go ahead and uh print that.

There we go.

And then we'll go ahead and do the uh np

dot of u comma

v.

And we'll find when we do this, we go

and run this uh we're going to get uh

370

370.

So this is a strain multiplication where

they use it to solve uh linear algebra

uh when you have multiple numbers going

across. And so this could be very

complicated. We could have a whole

string of different variables going in

here. But for this we get a nice uh

value for our dot multiplication

and we did um addition earlier which was

just your basic addition. Uh and of

course a matrix you can get very

complicated on these or in this case

we'll go ahead and do um let's create

two complex matrixes.

This one is a matrix of um you know 1210

46 431. We'll just print out A so you

can see what that looks like. Here's

print A.

We print A out. You can see that we have

a um 2x3

layer matrix for A. And we can also put

together always kind of fun when you're

playing with print values. Uh we could

do something like this. We could go in

here. There we go. Uh, we could print a.

We have it end with uh equals a run. And

this kind of gives it a nice look. Uh,

here's your matrix. That's all this is.

Comma, n means it just tags it on the

end. That's all all that is doing on

there. And then we can simply add in

what is a plus b. And you should already

guess because this is the same as what

we did before. There's no difference.

Uh, we do a simple vector addition. We

have 12 + 2 is 14, 10 + 8 is 18. And so

on. And just like we did the uh matrix

addition, we can also do a minus b and

do our matrix subtraction.

And we look at this uh we have what? 12

- 2 is 10. 10 - 8 um where are we?

Oh, there we go. 8 min

confusing what I'm looking at. I should

have reprinted out the original numbers.

Uh but we can see here 12 - 2 is of

course 10. 10 - 8 is 2. Uh 4 - 46 is -

42 and so forth. So same as a

subtraction as before, we just call it

matrix subtraction. It's identical.

Now if you remember up here, we had

scalar addition where we're adding just

one number to a matrix. You can also do

scalar multiplication. Uh and so simply

if you have a single value A and you

have B which is your array, we can also

do A * B. When we run that, uh, you can

see here we have 2 * 4 is 8. Uh, 5 * 4

is 20 and so forth. You're just

multiplying the four across each one of

these values. And this is an interesting

one that comes up. A little bit of a

brain teaser is matrix and vector

multiplication.

And so when we're looking at this,

uh, we are just do a regular arrays. It

doesn't necessarily have to be a numpy

array. We have a

which has our um array of arrays and b

which is a single array and so we can

from here

do the dot

a b and this is going to return two

values and the first value is that it's

you could say it's like uh um we're

doing the this array b array first with

a and then with a second one and so it

splits it up so you have a matrix of

vector multiplication and you mix and

match. When you get into really

complicated uh backend stuff, this

becomes more common because you're now

you got layers upon layers of data and

so you you'll end up with a matrix and a

set of uh vector matrices. Do you want

to multiply?

Now, keep in mind that if you're doing

data science, a lot of times you're not

looking at this. This is what's going on

behind the scenes. So if you're in um

the scikit looking at sklearn where

you're doing linear regression models,

this is some of the math that's hidden

behind the scenes that's going on. Other

times you might find yourself having to

do part of this and manipulate the data

around so it fits right and then you go

back in and you run it through the

scikit. And if we can do um up here

where we did a uh matrix and vector

multiplication, we can also do matrix to

matrix multiplication. And if we run

this where we have the two matrices, uh

you can see we have a very complicated

array that of course comes out on there

for our dot. And just to reiterate it,

we have our transpose a matrix which is

your T. And so if we create a matrix A

and then we do transpose it, you can see

how it flips it from 5 10 15 20 25 30 to

5 15 25 10 20 30 uh rows and columns.

And certainly with the math, uh, this

comes up a lot. Um, it also comes up a

lot with XY plotting. When you put it

into piplot, you have one format where

they're looking at pairs of numbers and

then they want all of X's and all Y's.

So, you know, the transpose is an

important tool both for your math and

for plotting and all kinds of things.

Another tool that we didn't discuss uh

is your identity matrix. Uh and this one

is more definition.

Uh the identity matrix. Um we have here

one where we just did uh two. So it

comes down as one 0 0 1 uh 1 0 0 1 0. It

creates a diagonal of one. And what that

is is when you're doing your identities,

you could be comparing all your

different features to the different

features and how they correlate. And of

course when you have uh feature one

compared to feature one to itself it is

always one uh where usually it's between

zero one depending on how well

correlates. So when we're talking about

identity matrix that's what we're

talking about right here is that you

create this preset matrix and then you

might adjust these numbers depending on

what you're working with and what the

domain is. And then another thing we can

do uh to kind of wrap this up. We'll hit

you with the most complicated uh um

piece of this puzzle here is an inverse

um a matrix. And let's just go ahead and

put the um it's a lengthy description.

Let's go and put the description. This

is straight out of the uh the website

for um numpy. Uh so given a square

matrix A, here's our square matrix A,

which is 2 1 0 0 1 0 1 2 1. Keep in mind

3x3, it's square. It's got to be equal.

It's going to return the matrix A

inverse satisfying dot A um A inverse.

So here's our matrix multiplication.

Um and then of course it equals the dot

uh yeah a inverse of a um with an

identity shape of uh a dotshaped zero.

This is just reshaping the identity.

That's a little complicated there. Uh so

we go and have our here's our array. Uh

we'll go ahead and run this. And you can

see what we end up with is we end up

with uh an array 0.5 minus 0.5 and so

forth with our 211 going down to 1 0 0 1

0 1 2 1. Um getting into a little deep

on the math understanding when you need

this is probably really is is what's

really important when you're doing data

science versus uh handwriting this out

and looking up the math and handwriting

all the pieces out. you do need to know

about the linear algorithm inverse of a.

Uh so if it comes up, you can easily

pull it up or at least remember where to

look it up. You took a look at the

algebra side of it. Let's go ahead and

take a look at the calculus side of uh

what's going on here with the machine

learning. So calculus, oh my goodness,

and differential equations, you got to

throw that in there because that's all

part of the bag of tricks, especially

when you're doing large neural networks,

but also comes up in many other areas.

The good news is most of it's already

done for you in the back end. Uh so when

it comes up, you really do need to

understand from the data science, not

data analytics. Data analytics means

you're digging deep into actually

solving these math equations. U and a

neural network is just a giant

differential equation. Uh so we talk

about calculus uh we're going to go

ahead and understand it by talking about

cars versus time and speed. uh so helps

to calculate the spontaneous rate of

change.

Uh so suppose we plot a graph of the

speed of a car with respect to time. So

as you can see here going down the

highway probably merged into the highway

from an on-ramp. So I had to accelerate

so my speed went way up uh stuck in

traffic merged into the traffic. Traffic

opens up and I accelerate again up to

the speed limit and u maybe it peters

off up there. So you can look at this as

as um the speed versus time. I'm getting

faster and faster because I'm

continually accelerating. And if I hit

the brakes, it go the other way. So the

rate of change of speed with respect of

time is nothing but acceleration. How

fast are we accelerating? The

acceleration is the area between the

start point of x and the end point of

delta x. Uh so we can calculate a simple

if you had x and delta x we could put a

line there and that slope of the line is

our acceleration.

Now that's pretty easy when you're doing

linear algebra but I don't want to know

it just for that line and those two

points. I want to know it across the

whole of what I'm working with. That's

where we get into calculus. So when we

talk about the distance between x and

delta x it has to be the smallest

possible near to zero in order to

approximate the acceleration.

Uh so the idea is that instead of I mean

if you ever did took a basic calculus

class they would draw bars down here and

you would divide this area up um let's

go back up a screen. you divide this

area of this time period up into maybe

10 sections and you'd use that and you

could calculate the acceleration between

each one of those 10 sections kind of

thing. Uh and then we just keep making

that space smaller and smaller until

delta x is almost uh infantismally

small. And so we get a function of a uh

equals a limit as h goes to zero of a

function of a plus h minus a function of

a over h. And that is you're computing

the slope of the line.

We're just computing that slope under

smaller and smaller and smaller samples.

Uh and that's what calculus is. Calculus

is the integral. You can see down here

we have our nice uh integral sign. Looks

like a giant s. And that's what that

means is that we've taken this down to

as small as we can for that sampling. Uh

so we're talking about calculus. Finding

the area under the slope is the main

process in the integration. Similar

small intervals are made of the smallest

possible length of x plus delta x where

delta x approaches almost an infantismly

small space. And then it helps to find

the overall acceleration by summing up

all the lengths together. Uh so we're

summing up all the accelerations from

the beginning to the end. And so here's

our integral. we sum of a of x * d ofx =

a + c. Uh that is our basic calculus

here. So when we talk about

multivvariant calculus, uh multivariate

calculus deals with functions that have

multiple variables and you can see here

we start getting into some very

complicated equations. Um uh change in w

over change of time equals change of w

over change of z. the differential of z

to dx differential of x to dt. It gets

pretty complicated. Uh and it really

translates into the multivariate

integration using double integrals. And

so you have the the sum of the sum of f

ofxy of d of a equals the sum from c to

d and a to b of f ofxy dx dy equals uh

the sum of a to b sum of c to d of fxy

dy dx.

understanding the very specifics of

everything going on in here and actually

doing the math is usually calculus one,

calculus 2, and differential equations.

Uh so you're talking about three

fulllength courses to dig into and solve

these math equations. What we want to

take from here is we're talking about

calculus. Uh we're talking about summing

of all these different slopes. And so

we're still solving a linear uh

expression. We're still solving y = mx +

b, but we're doing this for

infantismally small x's. And then we

want to sum them up. That's what this

integral sign means. The the sum of a of

x d of x= a plus c.

And when you see these very complicated

uh multivariate differentiation using

the chain rule uh when we come in here

and we have the change of w to the

change of t equals the change of w dz uh

and so forth. That's what's going on

here. That's what these means. We're

basically looking for the area under the

curve which really comes to how is the

change changing and speed's going up.

How is that changing? And then you end

up with a multiple layer. So if I have

three layers of neural networks, how is

the third layer changing based on the

second layer changing which is based on

the first layer changing? And you get

the picture here that now we have a very

complicated uh multivariate integration

um with integrals.

The good news is we can solve this uh

mathematically and that's what we do

when you do neural networks and reverse

propagation. Uh so the nice thing is

that you don't have to solve this on

paper unless you're a data analysis and

you're working on the back end of

integrating these formulas and building

the script to actually build them. So we

talk about applications of calculus. Uh

it provides us the tools to build an

accurate predictive model. Um so it's

really behind the scenes we want to

guess at what the change of the change

of the change is.

That's a little goofy. I I know I just

threw that out there. It's kind of a

meta term. But if you can guess how

things are going to change, then you can

guess what the new numbers are.

Multivariate calculus explains the

change in our target variable in

relation to the rate of change in the

input variables. So there's our multiple

variables going in there. If uh one

variable is changing, how does it affect

the other variable? And then in gradient

descent, calculus is used to find the

local and global maxima. And this is

really big. Uh we're actually going to

have a whole section here on gradient

descent because it is really I mean I

talked about neural networks and how you

can see how the different layers go in

there, but gradient descent is one of

the most key things for trying to guess

the best answer to something. So let's

take a look at the code behind gradient

descent. And uh before we open up the

code, let's just do real quick uh

gradient descent.

Let's say we have a curve like this. And

most common is that this is going to

represent your error. Oops.

Error. There we go. Error. Ah, hard to

read there. And I want to make the error

as low as possible. And so what I'm

looking at it is I want to find this

line here which is the minimum value. So

we're looking for the minimum and it

does that by uh sampling there and then

it based on this it guesses it might be

someplace here and it goes hey this is

still going down. It goes here and then

goes back over here and then goes a

little bit closer and it's just playing

a high low until it gets to that spot,

that bottom spot. And so we want to

minimize the error in uh on the flip

note, you could also want to be

maximizing something. You want to get

the best output of it. Uh that's simply

uh minus the value. Uh so if you're

looking for where the peak is, this is

the same as a negative for where the

valley is and looking for that valley.

Uh that's all that is and this is a way

of finding it. So the cool thing is um

all the heavy lifting's done. Um I

actually ended up putting together one

of these a while back as uh when I

didn't know about sidekick and I was

just starting. Boy, it's a long while

back and uh is playing high low. How do

you play high low? not get stuck in the

valleys, uh, figure out these curves and

things like that. Well, you do that and

the back end is all the calculus and

differential equations to calculate this

out. The good news is you don't have to

do those. Uh, so instead, we're going to

put together the code and let's go ahead

and see what we can do with that.

So, uh, guys in the back put together a

nice little piece of code here, which is

kind of fun.

uh some things we're going to note and

this is this is really important stuff

because when you start doing your data

science and digging into your machine

learning models uh you're going to find

these things are stumbling blocks. Uh

the first one is current x. Where do we

start at? Uh keep in mind your model

that you're working with is very

generic. So whatever you use to minimize

it the first question is where do we

start? Um, and we started at this cuz

the algorithm starts at x= 3. So, we

arbitrarily picked five. Learning rate

is uh how many bars to skip going one

way or the other. Uh, I'm in fact, I'm

going to separate that a little bit

because these two are really important.

Um, if we're dealing with something like

this where we're talking about um uh

well, here's our here's the function

we're going to use our um gradient of

our function um 2 * x + 5. Keep it

simple. So that's a function we're going

to work with. So if I'm dealing with

increments of a th00and 0.1 is going to

be a very long time. And if I'm dealing

with increments of 0.001,

uh 0.1 is going to skip over my answer.

So I won't get a very good answer. Um

and then we look at precision. This

tells us when to stop the algorithm. So

again, very specific to what you're

working on. uh if you're working with

money and you don't convert it into a

float value uh you might be dealing with

0.01 which is a penny that might be your

precision you're working with. Um and

then of course the previous step size

max iterations uh we want something to

cut out at a certain point. Usually

that's built into a lot of minimization

functions. And then here's our actual uh

formula we're going to be working with.

And then we come in, we go while

previous step size is greater than

precision and its is less than max its

say that 10 times fast. Um

we're just saying if it's uh if we're if

we're still greater than our precision

level, we still got to keep digging

deeper. Um and then we also don't want

to go past a thou or whatever this is, a

million or 10,000 uh running. That's

actually pretty high. um almost never do

max iterations more than like 100 or

200. Rare occasions you might go up to

four or 500 if it's depending on the

problem you're working with. Uh so we

have our previous equals our current.

That way we can track timewise.

Uh the current now equals the current

minus the rate times the formula of our

previous x. So now we've generated our

new version. Uh previous step size

equals the absolute current previous.

Uh, so we're looking for the change in x

itters equals iterations + one. That's

so we know to stop if we get too far.

And then we're just going to print the

local minimum occurs at x on here. And

if we go ahead and run this,

uh, you can see right here it gets down

to this point and it says, hey, um,

local minimum is minus 3.3222

for this particular series we created.

Uh, and this is created off of our

formula here. lambda x2 * x + 5. Now,

when I'm running this stuff, uh you'll

see this come up a lot

and uh with the sklearn kit and and one

of the nice reasons of breaking this

down the way we did is I could go over

those top pieces. Uh those top pieces

are everything when you start looking at

these minimization toolkits in built-in

code. And so from um we'll just do it's

actually docs.cipi.org

and we're looking at the scikit. There

we go. Um optimize minimize. You can

only minimize one value. You have the

function that's going in. This function

can be very complicated. Uh so we used a

very simple function up here. It could

be there's all kinds of things that

could be on there. And there's a number

of methods to solve this as far as how

they shrink down. Uh and your x knot.

There's your there's your start value.

So your function, your start value. Um

there's all kinds of things that come in

here that we can look at which we're not

going to. Um optimization automatically

creates constraints bounds. Some of this

it does automatically, but you really

the big thing I want to point out here

is you need to have a starting point.

You want to start with something that

you already know is mostly the answer.

Uh if you don't, then it's going to have

a heck of a time trying to calculate it

out.

Or you can write your own little script

that does this and and does a high low

guessing and tries to find the max

value. That brings us to statistics.

What this is kind of all about is

figuring things out. Lot of vocabulary

and statistics. Uh so statistics, well,

I guess it's all relative. It's

definitely not an ed class. Uh so a

bunch of stuff going on. Statistics.

Statistics concerns with the collection,

organization, analysis, interpretation,

and presentation of data. That is a

mouthful. Um so we have from end to end

we're

valid, what does it mean? How do we

organize it? Um how do we analyze it?

Then you got to take those analysis and

interpret it into something that uh

people can use. kind of reduce it to

understandable. Um, and nowadays you

have to be able to present it. If you

can't present it, then no one else is

going to understand what the heck you

did.

So, we look at the terminologies. Uh,

there is a lot of terminologies

depending on what domain you're working

in. So clearly if you're working in um a

domain that deals with

viruses and tea cells and and how does

you know where does that come from and

you're studying the different people

then you're going to have a population.

if you are working with um mechanical

gear um you know a little bit different

if you're looking for the wobbling

statistics uh to know when to replace a

rotor on a machine or something like

that uh that can be a big deal. You

know, we have these huge fans that turn

in our sewage processing systems. And so

those fans, they start to wobble and hum

and do different things that the sensors

pick up. At one point, do you replace

them? Instead of waiting for it to

break, in which case it cost a lot of

money. Instead of replacing a bushing,

you're replacing the whole fan unit. Uh

an interesting project that came up for

our city a while back. Uh so population,

all objects are measurements whose

properties are being observed. Uh so

that's your population all the objects.

It's easy to see it with people because

we have our population and large. Um but

in the case of the sewer fans we're

talking about how the fan units. That's

the population of fans that we're

working with.

You have a parameter a matrix uh that is

used to represent a population or

characteristic.

You have your sample a subset of the

population studied. You don't want to do

them all because then you don't have a

if you come up with a conclusion for

everyone, you don't have a way of

testing it. So you take a sample. Uh

sometimes you don't have a choice. You

can only take a sample of what's going

on. You can't u study the whole

population. And a variable, a metric of

interest for each person or object in a

population.

Types of sampling. We have probabilistic

approach. uh selecting samples from a

larger population using a method based

on the theory of probability

and we'll go into a little bit more

deeper on these. We have random

systematic stratified and then you have

nonprobabilistic approach selecting

samples based on the subjective judgment

of the researcher rather than random

selection. Uh it has to do with

convenience trying to reach a quota um

or snowball. Uh and they're very biased.

That's one of the reasons you'll see

this big stamp on it says biased. Uh so

you got to be very careful on that. So

probabilistic sampling uh when we talk

about a random sampling, we select

random size samples from each group or

category. So we it's as random as you

can get. Uh we talk about systematic

sampling. We're selecting randomsiz

samples from each group or category with

a fixed periodic interval. Uh so we kind

of split it up. This would be like a

time setup or different categories. And

you might ask your question, what is a

category or a group? Uh if you look at

I'm going to go back a window. Let's say

we're studying um economics of different

of an area. Um we know pretty much that

based on their culture, where they came

from, they might need to be separated.

And so uh and when I say separated, I

don't mean separated from their their uh

place where they live. I mean, as far as

the analysis, we want to look at the

different groups and make sure they're

all represented. So, if we had like an

80% uh of a group that is uh say

Hispanic and or Indian and also in that

same area, we have 20% 20% who are let's

call our expatriots. They left America

and they're nice and uh your Caucasian

group. We might want to sample a group

that is representative of both. Uh, so

we're talking about stratified sampling

and we're talking about groups. Those

are the groups we're talking about. And

it brings us to stratified sampling,

selecting approximately equalized

samples from each group or category. Uh,

this way we can actually separate the

categories and give us an insight into

the different cultures and how that

might affect them in that area. Uh so

you can see these are very very

different kind of depends on what you're

working with um as far as your data and

what you're studying. And so we can see

here just to go a little bit more we'd

have selecting 25 employees from a

company of 250 employees randomly. Don't

care anything about them. What groups

they're in, which office are in,

nothing. Um and we might be selecting

one employee from every 50 unique

employees in a company of 250 employees.

And then we have selecting one employee

from every branch in the company office.

So we have all the different branches.

There's our group or our categories by

the branch. And the category could

depend on what you're studying. So it

has a lot of variation on there. You see

this kind of grouping and categorizing

is also used to generate a lot of

misinformation.

Uh so if you only study one group and

you say this is what it is, then

everybody assumes that's what it is for

everybody. And so you got to be very

careful of that. and it's very unethical

thing to kind of do. So, types of

statistics. Uh we talk about statistics,

we're going to talk about descriptive

and inferential statistics. There are so

many different terms in statistics to

break it up. Uh so we so we're talking

about a particular setup. So we're

talking about descriptive and

inferential uh statistics. You the base

of the word describe is pretty solid.

you're describing the data. What does it

look like? With inferial statistics,

we're going to take that from the small

population to a large population. So, if

you're working with a drug company, uh

you might look at the data and say,

"These people were helped by this drug.

They did uh 80% better as far as their

health or 80% better survival rate than

the people um who did not have the drug.

So, we can infer that that drug will

work in the greater populace and will

help people. So that's where you get

your inferential. Uh so we are

predicting how it's going to affect the

greater population.

So descriptive statistics it is used to

describe the basic features of data and

form the basis of quantitative analysis

of data. So we have a measure of central

tendencies. We have your mean, median

and mode. And then we have a measure of

spread like your range, your

interquartile range, your variance and

your standard deviation. And we're going

to look at all these a little deeper

here in a second. Uh but one of them you

can think of is um how the data

difference differences you know what's

the max men range all that stuff is your

spread and anything that's just a single

number is usually your central uh

tendencies measure of central

tendencies. So we talk about the mean it

is the average of the set of values

considered. what is the average outcome

of whatever's going on? And then your

median separates the higher half and the

lower half of data.

Uh so where's the center point of all

your different data points? So your mean

might have some a couple really big

numbers that skew it uh so that the

average is much higher than if you took

those outliers out where the median

would by separating the high from the

low might give you a much lower number.

you might look at and say, "Oh, that's

that's odd. Why is the average so much

higher than the median?" Well, it's

because you have some outliers, or why

is it so much lower? And then the mode

is the most frequent appearing value.

Uh, this is really interesting. If

you're studying economics and how people

are doing, you might find that the most

common um income like in the US was at

one point 24,000 a year where the

average was closer to 80,000. And it's

like, wow, what a difference. Well,

there's some people have a lot of money

and so that skews that way up. So the

average person is not making that kind

of money. And then you look at the

median income and you're like, well, the

median income is a little bit closer to

the average. Uh so it does create a very

interesting way of looking at the data.

Again, these are all uh central

tendencies, single numbers you can look

at for the whole spread of the data.

And we look at the measure of central

tendencies. The mean is the average

marks of a students in a classroom. So

here we have the mean sum of the marks

of the students total number of students

and as we talked about the median uh if

we have 0 through 10 and we take half

the numbers and put them on one side of

the line half the numbers on the other

side of the line uh we end up with five

in the middle and then the mode what

mark was scored by most of the students

in a test in a simple case where most

people scored like an 82% and got

certain problems wrong easy to figure

out. uh not so easy when you have

different areas where like you have like

the um oh let's go back to economy a

little bit more difficult to calculate

if you have a large group that scores

that makes 30,000 and a slightly bigger

group that makes 26,000. So what do you

put down for the mode? Uh certainly

there's a number of ways to calculate

that and there's actually a different

variations depending on what you're

doing. So now we're looking at a measure

of spread uh range. What's the

difference between the highest and the

lowest value? First thing you want to

look at, you know, it's we had everybody

in the test scored between 60 and 100%,

somebody got 100% or maybe 60 to 90%. It

was so hard that a lot of people could

not get 100%. Um, and you have your

interquartile range. Quartortiles divide

a rankorder data set into four equal

parts.

very common thing to do as part of all

the basic packages whether you're

working in uh dataf frames with pandas

whether you're working in scala whether

you're working in R um you'll see this

come up where they have range your min

your max and then it'll have your

interquartile range how does it look

like in each quarter of data variance

measures how far each number in the set

is from the mean and therefore from

every other number in the set uh so you

have like a how much turbulence is going

on in this data. And then the standard

deviation, it is the measure of the

variance or the dispersion of a set of

values from the mean. And you'll usually

see uh if I'm doing a graph, I might

have the value graphed. Um and then

based on the the error, I might graph

graph the standard deviation and the

error on the graph as a background so

you can see how far off it is. Uh so

standard deviation is used a lot. So

measurement of spread uh marks of a

student out of a 100 uh we have here

from 50 to 63 or 50 to 90 uh so the

range maximum marks minimum marks we

have 90 to 45 and the spread of that is

45 90 - 45 and then we have the

interquartile range using the same marks

over there you can see here where the

median is and then there's the first

quarter the second quarter and the third

quarter based on splitting it apart by

those values

And to understand the variance and

standard deviation, we first need to

find out the mean. Uh so here's our our

you know calculating the average there.

We end up at approximately 66 for the

average. And then we look at that the

variance once we know the means we can

do equals the marks minus the mean

squared. Why is it squared? Uh because

one, you want to make sure it's you

don't have like if you if you're putting

all this stuff together, you end up with

an error as far as one's negative, one's

positive, one's a little higher, one's a

little lower. Uh so you always see the

squared value and over the total

observations. And so the standard

deviation equals the square root of the

variance, which is approximately 16. And

if you were looking at um a predictable

model, you would be looking at the

deviation based on the error. How much

error does it have? Uh that's again

really important to know if you're if

your prediction is predicting something,

what's a chance of it being way off or

just a little bit off.

Now that we've looked at the um tools as

far as some of the basics for doing your

statistics and what we're talking about,

let's go ahead and pull up a little demo

and show you what that looks like in

Python code. Uh so you can get some

little hands-on here. For that, let's go

back into our Jupyter notebook in

Python. Now, almost all of this you can

do in numpy. Last time we worked um in

numpy. This time we're going to go ahead

and use pandas. And if you remember from

pandas on here, uh this is basically a

data frame, rows, columns. Let's just go

ahead and do a print df. head

and run that.

And you can see we have uh the name

Jane, Michael, William, Rosie, Hannah,

and their salaries on here. And of

course, instead of having to do all

those hand calculations and add

everything together and divide by the

total, we can do something very simple

on this uh like use the command mean in

pandas. And so if I go ahead and do this

print df, pick our column salary because

we want to find the means of that

colery.

We want to find the means of that

column. Uh and we go and print this out.

And you can see that the uh average

income on here is 71,000.

Uh, and let's just go ahead and do this.

We'll go ahead and put in uh means.

And if we're going to do that, we also

might want to find the median.

And the median is uh very similar except

it actually is just median. Uh we're

used to means and average. It's kind of

interesting that those are they use the

two different words. Uh there can be in

some computations slight differences but

for the most part the means is the

average. Uh and then the median oops

let's put a

median here. DF salary that way it

displays a little better. We can see the

median is 54 um000. So the halfway mark

is significantly below the average. Why?

Because we have somebody in here who

makes 189,000.

Darn you Rosie for throwing off our

numbers. Uh but that's something you'd

want to notice. This is this is the

difference between these is huge and so

is what is the meaning behind that when

you're studying a populace and looking

at uh the different data coming in. And

of course we also want to find out hey

what's the most uh common income that

people make in this little tiny sample.

And so we'll go ahead and do the mode.

And you can see here with the mode uh

it's at 50,000.

So this is this is very telling that

most people are making 50,000. The

middle point is at 54,000. So half the

people are making more than that. What

that tells me is that if the most common

income is way is below the median, then

there's a few there's a SK, you know,

there's a a lot of high salaries going

up, but there's some really low salaries

in there. And so this trend which is

very common in statistic you when you're

analyzing the economy and different

people's income is pretty common and the

bigger difference between these is also

very important when we're studying

statistics. Uh and when you hear someone

just say hey the average income was you

might start asking questions at that

point. Why aren't you talking about the

median income? Why aren't you talking

about the mode the most common income?

What are you hiding? Uh and if you're

doing these analysis, you should be

looking at these saying, "Hey, why why

are this discrepancies? Why are these so

different?" And of course, with any uh

analysis, it's important to find out the

minimum

and the maximum. So, we'll go ahead.

It's just simply uh um min'll pull up

your minimum and then do max pulls up

the maximum. pretty straightforward on

as far as um translating it and knowing

what your you know what the your lowest

value and what your highest value is

here. Um which you'll use to generate

like a spread later on. And real quick

on no mode mode, uh note that it puts

mode zero. Like I said, there's a couple

different ways you can compute the mode.

Um although, you know, standard one's

pretty good. We can of course do the

range, which is your max minus your min.

So now we have a range of 149,000

between the upper end and the lower end.

And you might want to be looking up the

individual values on all of these. But

it turns out there is a describe

feature in pandas.

And so in pandas we can actually do df

salary describe. And if we do this you

can see we have that there's seven uh

setups. Here's our mean. Um, our

standard deviation, which we didn't

compute yet, which would just be a STD.

And you got to be a little careful

because when it computes it, it looks

for axes and things like that. Uh, we

have our minimum value, and here's our

cortiles,

uh, our maximum value, and then of

course the name salary. Uh, so these are

the these are the basic statistics. You

can pull them up and just describe. This

is a dictionary. So I could actually do

something like um in here I could

actually go uh count and run. And now it

just prints the count. Uh so because

this is a dictionary, you can pull any

one of these values out of here. It's

kind of a quick and dirty way to pull

all the different information and then

split it up and depending on what you

need. Now if I just walked in and gave

you this information um in a meeting, at

some point you would just kind of fall

asleep. That's what I would do anyway.

Um, so we want to go ahead and and see

about graphing it here. And we'll go

ahead and put it into a histogram and

plot that graph on it of the salaries.

And let's just go ahead and put that in

here. So we do our map plot inline.

Remember that's a Jupiter's notebook

thing. Uh, a lot of the new version of

the mapplot library does it

automatically, but just in case I always

put it in there. Uh, import mattplot

library piplot as plt. That's my

plotting.

And then we have our data frame. Uh I

don't I guess I really don't need to

respell the data frame. Maybe we could

just remind oursel what's in it. So

we'll go ahead and just uh print

DF. That way we still have it. And then

we have our salary. DF salary

salary.plot history title salary

distribution color gray. Uh plot AXV

line salary the mean value. So, we're

going to take the mean value um color

violet line style dash. This is just all

making it pretty. Uh what color dash

line width of two that kind of thing.

And the median. And let's go ahead and

run this just so you can see what we're

talking about.

And so up here we are taking on our

plot. Um so here's the data. Here's our

our data frame printed out so you can

see it with the salaries. We're looking

at the salary distribution and just look

at this the way they're the salary is

distributed. Um you have our in this

case we did let's see we had red for the

median we have violet

for our average or mean and you can just

see how it really here's our outlier.

Here's our person who makes a lot of

money. Here's the um average and here's

the median. Um, and so as you look at

this, you can say, "Wow." Um, based on

the average, it really doesn't tell you

much about what people are really taking

home. All it does is tell you how much

money is in this, you know, what the

average salary is. So, some of the

things you want to take away in addition

to this is that it's very easy to plot

um an AXV line. These are these up and

down lines for your markers. Um, and as

you display display the data, I mean,

you can add all kinds of things to this

and get really complicated. Keeping it

simple is pretty straightforward. I look

at this and I can see we have a major

outlier out here. We can definitely do a

histogram and stuff like that. Um, but

you know, picture's worth a thousand

words. What you really want to make sure

you take away is that we can do a basic

describe which pulls all this

information out and we can print any of

the individual information from the

describe uh because this is a

dictionary.

And so if we want to go ahead and look

up um the mean value, we can also do

describe mean. So if you're doing a lot

of statistics, uh being able to

doesn't have the print on there, so it's

only going to print um the last one,

which happens to be the mean. Uh you can

very easily reference any one of these.

And then you can also, if you're doing

something a little bit more complicated

and you don't need just the basics, you

can come through and pull any one of the

individual um

references from the from the pandas on

here. So now we've had a chance to

describe our data. Uh let's get into

inferential statistics. Inferial

statistics allows you to make

predictions or inferences from data. And

you can see here we have a nice little

picture movie ratings and um if we took

this group of people and said hey how

many people like the movie dislike it

can't say and then you ask just a random

person who comes out of the movie who

hasn't been in this study uh you can

infer that 55% chance of saying liked

35% chance of saying disliked or a 10 or

11% chance of can't say. So that that's

real basics of what we're talking about

is you're going to infer that the next

person is going to follow these

statistics.

Uh so let's look at point estimation. Uh

it is a process of finding an

approximate value for a population's

parameter like mean or average from

random samples of the population. Let's

take an example of testing vaccines for

COVID 19. Uh vaccines and flu bugs, all

that. It's a pretty big thing of how do

you test these out and make sure they're

going to work on the populace. A group

of people are chosen from the

population. Medical trials are

performed. Results are generalized for

the whole population. So here's a

protected here's our small group up here

where we've selected them. We run

medical trials on them and then the

results work for the population. You

nice diagram with the arrows going back

and forth and the very scary co virus in

the middle of one. And let's take a look

at the applications of inferial

statistics.

Very central is what they call

hypothesis testing uh and the confidence

interval which go with that. And then as

we get into

probability, we get into our binomial

theorem, our normal distribution and

central limit theorem. Hypothesis

testing. Hypothesis testing is used to

measure the plausibility of a hypothesis

assumption by using sample data. Now

when we talk about theorems, theory,

hypothesis,

uh keep in mind that if you are in a

philosophy class, theory is the same as

hypothesis where theorem is a scientific

uh statement that is something that has

been proven although it is always up for

debate because in science we always want

to make sure things are up to debate. So

a hypothesis is the same as a phil

philosophical class calling a theory

where theory in science is not the same.

Theory in science says this has been

well proven. Gravity is a theory. Uh so

if you want to debate the theory of

gravity try jumping up and down. If you

want to have a theory about why the

economy is collap collapsing in your

area that is a philosophical debate.

Very important. I've heard people mix

those up and it is a pet peeve of mine.

When we talk about hypothesis testing,

the steps involved in hypothesis testing

is first we formulate a hypothesis. We

figure out the right test to test our

hypothesis. We execute the test and we

make a decision. And so when you're

talking about hypothesis, you're usually

trying to disprove it. If you can't

disprove it and it works for all the

facts, then you might call that a

theorem at some point. So in a use case,

uh let's consider an example. We have

four students. were given a task to

clean a room every day. Sounds like

working with my kids. They decided to

distribute the job of cleaning the room

among themselves. They did so by making

four chits which has their names on it

and the name that gets picked up has to

do the cleaning for that day. Rob took

the opportunity to make chits and wrote

everyone's name on it. So here's our

four people, Nick, Rob, Imlia, Imlia,

and Summer.

Now Rick, Imlia and Summer are asking us

to decide whether Rob has done some

mischief in preparing the chits i.e

whether Rob has written his name on one

of the chit. For that we will find out

the probability of Rob getting the

cleaning job on first day, second day,

third day and so on till 12 days. The

probability of Rob getting the job

decreases every day. I.e. his turn never

comes up. Then definitely he has done

some mischief while making the chits. So

the probability of Rob not doing work on

day one is uh three out of four. There's

a 75 chance that he didn't do work. Uh

two days 34s * 34s equals.56.

3 days you have 3/4 34 34 which

equals42.

Uh when you get to day 12 it's 0032

which is less than 0.05.

Remember this 0.05 uh that comes up a

lot when we're talking about um certain

values when we're looking at statistics.

Rob is cheating as he wasn't chosen for

12 consecutive days. That's a very high

probability when on day 12 he still

hasn't gotten the job cleaning the room.

So we come up to our important important

terminologies.

We have null hypothesis.

a general statement that states that

there is no relationship between two

measured phenomenon or no assoc

association among the groups.

Alternative hypothesis contrary to the

null hypothesis it states whenever

something is happening a new theory is

preferred instead of an old one. And so

the two hypothesis go hand in hand. Uh

so your null this is always interesting

in in we're talking about data science

and the math behind it. It's about

proving that the things have no

correlation. Null hypothesis says these

two have zero relation to each other.

Where the alternative hypothesis says,

hey, we found a relation. This is what

it is. We have p value. The p value is

the probability of finding the observed

or more extreme results when the null

hypothesis of a study question is true.

And the t value, it is simply the

calculated difference represented in

units of standard error. The greater the

magnitude of t, the greater the evidence

against the null hypothesis. And you can

look at the t value as being specific to

the test you're doing where the p value

is derived from your t value and you're

looking for what they call the 5% or the

0.05

showing that it has a high correlation.

So digging in deeper, let's assume that

a new drug is developed with the goal of

lowering the blood pressure more than

the existing drug. And this is a good

one because uh the null value here isn't

that you don't have any drug. The null

value here is that it's better than the

existing drug. The new drug doesn't

lower the blood pressure more than the

existing drug. Now if we get that uh

that says our null hypothesis is

correct. There is no correlation and the

new drug is not doing its job. The

alternative hypothesis the new drug does

significantly lower the blood pressure

more than the existing drug. Uh, yay, we

got a new drug out there. And that's our

alternative hypothesis or the H1 or HA.

And we look at the p value results from

the evidence like medical trials showing

positive results which will reject the

null hypothesis. And again, they're

looking for um a 0.05 or 5%. And the t

value comparing all the positive test

results and finding means of different

samples in order to test hypothesis. So

this is specific to the test. how uh

what percentage of increase did they

have and this leads us to the confidence

intervals. Uh a confidence interval is a

range of values we are sure our true

values of observations lie in. Let's say

you asked a dog owner around you and

asked them how many cans of food do you

buy for your uh per year for your dog.

Through calculations you got to know

that the on an average around 95% of the

people bought around 200 to 300 cans of

food. Hence we can say that we have a

confidence interval of 230 where 95% of

our values lie in that spread data

spread. Uh and this the graph really

helps a lot. So you can start seeing

what you're looking at here where you

have the 95%. You have your peak in this

case it's a normal distribution. So you

have the nice bell curve equal on both

sides. It's not asymmetrical. And 95% of

all the values lie within a very small

range. And then you have your outliers

the 2.5% going each way.

So we touched upon hypothesis uh and

we're going to move into probability. Uh

so you have your hypothesis. Once you've

generated your hypothesis, we want to

know the probability of something

occurring. Probability is a measure of

the likelihood of an event to occur. Any

event can be predicted with total

certainty and can only be predicted as a

likelihood of its occurrence. So any

event cannot be predicted with total

certainty. It can only be predicted as a

likelihood of its occurrence. Uh score

prediction. how good you're going to do

in whatever sport you're in, weather

prediction, stock prediction, if you've

studied physics and chaos theory, even

the location of the chair you're sitting

on has a probability that it might move

3 ft over. Granted, that probability is

one in like uh I think we calculated as

under one in trillions upon trillions.

So, it's the better the probability, the

more likely it's going to happen. There

are some things that have such a low

probability that we don't see them. So

we talk about a random variable. Uh

random variable is a variable whose

possible values are numerical outcomes

of a random phenomena. So uh we have the

coin toss. How many heads will occur in

the series of 20 coin flips? Probably

you know the on average there are 10,

but you really can't know because it's

very random. How many times a red ball

is picked from a bag of balls if there's

equal number of of red balls and blue

balls and green balls in there. How many

times the sum of digits on two dice uh

result or five each? Um so you know

there's how often you're going to roll

two fives on your pair of dice. So in a

use case uh let's consider the example

of rolling two dice. We have a random

variable outcome equals y. You can take

values 2 3 4 5 6 7 8 9 10 11 12. So we

have a random variable and a combination

of dice and instead of looking at how

many times um both dice were roll five

let's go ahead and look at a total sum

of five and you have in as far as your

random variables you can have a one four

equals 5 4 1 2 3 32 so four of those

roles can be four if you look at all the

different options you have four of those

random rolls can be a five and if we

look at the total number

which happens to be 36 different

options. Uh you can see that we have

four out of 36 chance every time you

roll the dice that you're going to roll

a total of five. You're going to have an

outcome of five. And uh we'll look a

little deeper as to what that means. Uh

but you could think of that at what

point if someone never rolls a five or

they always roll a five, can you say,

"Hey, that person's probably cheating."

uh we'll look a little closer at the

math behind that but let's just consider

this as one of the cases is rolling two

dice and gambling. There's also a

binomial distribution. It is the

probability of getting success or

failure as an outcome in an experiment

or trial that is repeated multiple

times. And the key is is by meaning two

binomial. Uh so passing or failing an

exam, winning or losing a game and

getting either head or tails. So if you

ever see binomial distribution, it's

based on a um true false kind of setup.

You win or lose. Let's consider a uh use

case and let's consider the game of

football between two clubs Barcelona and

Dortmund. The teams will have to play a

total of four matches and we have to

find out the chances of Barcelona

winning the series. So we look at the

total games and we're looking at five

different games or matches. Let's say

that the winning chance for Barcelona is

75% or 75. That means at each game they

have a 75% chance that they're going to

win that game and losing chances are 25%

or 0.25. Clearly 75 plus 0.25 equals 1.

So that accounts for 100% of the game.

Probability for getting K wins in n

matches is calculated.

And we we're talking like so if you have

five games uh and you want to know if I

play um how many wins in those five

games should I get? What's a percentage

on those? And the probability for

getting k wins in n matches is

calculated by px= k= n k p the k q to

the n minus k. Here p is the probability

of success and q is the probability of

failure. And so we can do total games of

n equals 5 where k equals 012345.

P which is the chance of winning is 75.

Q the chance of losing equals 1 minus p

which equals 1 - 0075 which equals 0.25.

The probability that Barcelona will lose

all of the matches can then just plug in

the numbers and we end up with a

09765625.

So very small chance they're going to

lose all their matches.

And we can plug in uh the value for two

matches. Probability that Barcelona will

win at least two matches is 00878. And

of course we can go on to probability

that Barcelona will win three matches

the 26 and of course four matches and so

on. And it's always nice to take this

information um and let let's find the

cumulative discrete probabilities for

each of the outcomes where Barcelona has

won three or more matches x= 3 x= 4 x= 5

and we end up with the p =264 plus 395 +

237 which equals89.

In reality the probability of Barcelona

winning the series is much higher than

75. And it's always nice to uh put out a

nice graph so you can actually see the

number of wins to the probability and

how that pans out with our binomial

case. Continuing in our important

terminology, location, the location of

the center of the graph depends on the

mean value. And uh this is some very

important things. So much of the data we

look at and when you start looking at

probabilities almost always has a

normalized look like the graph in the

middle.

uh but you do have left skewed where the

data is skewed off to the left and you

have more stuff happening off to the

left and you have right skewed data and

so when this comes up and these

probabilities come up where they're

skewed it's really important to take a

closer look at that uh mostly you end up

with a normalized set of data but you

got to also be aware that sometimes it's

a skewed data and then the height height

of the slope inversely depends upon the

standard deviation

so you can see down here the standard

deviation is really large it kind of

squishes it out. And if the standard

deviation is small, then most of your

data is going to hit right there in the

middle. You're going to have a nice

peak. Um, and so being aware of this

that you might have a probability that

fits certain data, but it has a lot of

outliers. So you're if you have a really

high standard deviation, um, if you're

doing stock market analysis,

this means your predictions are probably

not going to make you much money. uh

where if you have a very small

deviation, you might be right on target

and set to become a millionaire. Which

leads us to the zcore. Zcore tells you

how far from the mean a data point is.

It is measured in terms of standard

deviations from the mean. Around 68% of

the results are found between one

standard deviation. Around 95% of the

results are found between two standard

deviations.

And you read the symbols. Of course,

they love to throw some Greek letters in

there. we have mu minus 2 sigma. Mu is

just a quick way. It's that kind of

funky u. It just means the mean. Uh and

then the sigma is the standard

deviation. And that's the o with a

little arrow off to the right or the

little waggly tail going up. The o with

a with a line on it. Uh so mu minus 2

sigma is your uh 95% of the results are

found between two standard deviations.

The central limit theorem. This goes

back to the skew. If you remember, we

were looking at the skew values on this

previous slide. Have left skewed,

normalized, and right skewed. When we're

talking about it being skewed or not

skewed, the distribution of the sample

means will be approximately normally

distributed, evenly distributed, not

skewed. If you take large random samples

from the population with the mean mu and

the standard deviation sigma with

replacement

and you can see here um uh of course we

have our uh mu minus 2 sigma and the

spread down here the mean the median and

the mode and so when you're talking

about very large populations

these numbers should come together and

you shouldn't have a skewed value. If

you do that's a flag that something's

wrong. That's why this is so important

to be aware of what's going on with your

data, where your samples are coming

from, and the math behind it. And if

you're going to do all this, we got to

jump into conditional probability. The

conditional probability of an event A is

a probability that the event will occur

given the knowledge that an event B has

already occurred. And you'll see this as

Baze theorem. B A Y S bay. Uh, and this

is read. I mean, you have these funky

looking little P brackets. A B. This is

the probability of A being true while B

is already true. And you have the

probability of B being true when A is

already true. So, P B of A probability

of A being true divided by the

probability of B being true. And we talk

about BA's theorem which occurred back

in the 1800s when he discovered this.

This is such an important formula and

it's really it's not if you actually do

the math you could just kind of do um um

XY equals J K and then you divide them

out and you're going to see the same

math but it works with probabilities

which makes it really nice. And so if

you have a s you might have uh eight or

nine different studies going on in

different areas different people have

done the studies they brought them

together. Um if we look at today's co

virus the virus spread uh certainly the

studies done in China versus the studies

the way they're done in the US that data

is different in each of those studies

but if you can find a place where it

overlaps where they're studying the same

thing together you can then compute the

changes that you need to make in one

study to make them equal and this is

also true if you have a study of uh um

one group and you want to find out more

about it. So this formula is very

powerful. Uh it really has to do with

the data collection part of the math and

data science and understanding where

your data is coming from and how you're

going to combine different studies in

different groups. And we'll go ahead and

go into a use case. Uh let's find out

the chance of a person getting lung

disease due to smoking. Uh and this is

kind of interesting the way they word

this. Um let's say that according to

medical report provided by the hospital

states that around 10% of all patients

they treated suffered lung lung disease.

Uh so we have kind of a generic medical

report. They further found out uh by a

survey that 15% of the patients that

visit them smoke. So we have 10% that

are lung disease and um 15% of the

patients smoke. And finally, 5% of the

people continued smoke even when they

had lung disease. Uh not the brightest

choice um but you know it is an

addiction so it can be really difficult

to kick. And so we can look at the

probability of a uh prior probability of

10% people having lung disease. And then

probability b probability that a patient

smokes is 15%.

Uh and the probability of B um if B then

A. The probability of a patient smokes

even though they have lung disease is

5%. And probability of A is B.

Probability that the patient will have

lung disease if they smoke. And then

when you put the formulas together, uh

you get a nice solution here. You get

the probability of A of B, probability

that the patient will have lung disease

if they smoke. And you can just plug the

numbers right in and we get a 3.33%

chance. Hence, there is a 3.33% chance

that a person who smokes will get a lung

disease. So, we're going to pull up a

little Python code, always my favorite,

roll up the sleeves. Keep in mind, we're

going to be doing this um kind of like

the backend way so that you can see

what's going on. And then later on we're

going to create um we'll get into

another demo which shows you some of the

tools that are already pre-built for

this. Let's start by creating a set. So

we're going to create a set with curly

braces. This means that our set has um

only unique values. So you have a list

uh you have your tupils which can never

change and then you have um in this case

the the set. So 47, you can't create a

47, 4. It'll delete the four out. So

it's only unique values. And if you use

dictionaries,

quick reminder, this should look

familiar because it is a dictionary uh

where you have a value and that value is

assigned to or that key is assigned to a

value. Uh so you could have a key value

set up as a dictionary. So it's like a

dictionary without the value. It's just

the keys and they all have to be unique.

And if we run this, we have a set of 47.

We can also take a list, a regular um

setup. And I'm going to go ahead and

just throw in another number in here,

four, and run it. Uh, and you can see

here if I take my list 1 2 3 4, and I

convert it to a set, and here it is. My

set from list equals set my list.

The result is 1 2 3 4. So, it just

deletes that last four right out of

there.

And with the sets, you can also go in

there and um print here is my set. My

set uh three is in the set. And then if

you do three in my set,

that's going to be a logic function. Uh

and one in my set, six is not in the

set, and so forth. If we run this,

we get three is in the set true one is

in the set false because 357 is another

one. Six is in the set uh six is not in

the set. So not in my set. You can also

use this with a list. We could have just

used 357 and it would have um the same

response on there is three and usually

you do if three is in but three in my

set is still works on a just a regular

list. And we'll go ahead and do a little

iteration. We're going to do kind of the

dice one. Remember um uh 1 2 3 4 5 6.

And so we're going to bring in an

iteration tool and import product as

product.

And uh I'll show you what that means in

just a second. So we have our two dice.

We have dice A and it's going to be a

set of values. Um they can only have one

value for each one. That's why they put

it in a set. And if you remember from

range, it is up to seven. So this is

going to be 1 2 3 4 5 6. It will not

include the seven. And the same thing

for our dice B.

And then we're going to do is we're

going to create a list which is the

product of A and B. So what's um a + b?

And if we go ahead and run this uh it'll

print that out. And you'll see um in

this case when they say product because

it's an iteration tool,

we're talking about creating a tupole of

the two. So we've now created a tupole

of all possible outcomes of the dice

where dice A is one to three one to six

and dice B is 1 to six. And you can see

one to one, one to two, one to three and

so forth. You remember we had a slide on

this earlier where we talked about um

the different all the different outcomes

of a dice. We can play around with this

a little bit. Uh we can do in dice

equals two divi dice faces 1 2 3 4 5 6.

Uh another way of doing what we did

before and then we can create an event

space where we have a set which is the

product of the dice faces repeat equals

end dice. And we'll go ahead and just

run this. And you can see here it just

again puts it through all the different

possible variables we can have. And then

if we wanted to take the same uh set on

here and print them all out like we had

before uh we can just go through for

outcome and event space. Outcome end

equals. So the event space is creating

a sequence and as you can see here when

we print it out it stacks them versus

going through and putting them in a nice

line.

and we'll go ahead and do something. Um,

let's go print. Since we have the end

printing with a comma, that just means

it's just going to it's not going to hit

the return going down to the next line.

Uh, and we'll go ahead and do the length

of our event space. Uh, that'll be an

important variable we're going to want

to know in a minute.

And of course, if I get carried away

with my typing of length, uh, we'll

print it twice and it'll give me an

error. Uh so we have 36 different

possible variations here

and we might want to calculate something

like um what about the multiple of

three? What if we want to have

uh the probability of the multiple of

three in our setup?

And so uh we can put together the code

for the outcome in event space of xy

equals outcome if x + y

remainder 3. So, we're going to divide

by three and look at the remainder and

it equals zero.

Then it's a favorable outcome and we're

going to pop that outcome on the end

there.

And we'll turn it into a set. So, the

favor outcome equals a set. Not

necessary uh because we know it's not

going to be repeating itself, but just

in case, we'll go ahead and do that.

And if we want to print out the outcome,

we can go ahead and see what that looks

like. And you can see here these are all

uh multiples of three. Uh 1 plus 2 is 3,

5 + 4 is 9, which divided by 3 is 3, and

so forth.

And just like we looked up the length uh

of the one before, let's go ahead and

print the length of our f outcome so we

can see what that looks like.

There we go.

And of course, I did forget to add the

print in the middle because we're

looping through and putting an end on

the on the setup on there. So, we're

going to put the print in there. And if

I run this, you can see um

we end up with 12. So, we have 36 total

options. Uh we have 12 that are multiple

that um add up to a multiple of three.

And we can easily conver compute the

probability of this uh by simply taking

the length of our favorable outcome over

the length of the event space.

And if we print it out, let me put that

in there. Probability

last line. So we just type it in. We end

up with a 3333 chance. And it's roughly

a third.

And we might want to make this look

nice. So let's go ahead and put in

another line there. The probability of

getting the sum which is a multiple of

three is

3333.

We can compute the same thing for five

dice.

And if we do this for five dice and go

ahead and run it, you can see we just

have a huge amount of choices. So it

just goes on and on down here. And we

can look at the uh length of the event

space.

And we have over 7,776

choices. That's a lot of choices.

And if we want to ask the question like

we did above, uh what is the sum where

the sum is a multiple of five but not a

multiple of three? We can go through all

of these different options. And then uh

you can see here uh d1 d2 d3 d4 d5

equals the outcome. And if uh you add

these all together and the

division by five does not have a

remainder of zero but the remainder is

also of a division by three is not equal

to zero. So the multiple of five is

equal to zero but the multiple of three

is not. We can just appin that on here

and then we can look at that uh

favorable outcome. We'll go ahead and

set that and we'll just take a look at

this. What's our length of our favorable

outcome?

It's always good to see what we're

working with. And so we have 94 out of

776.

And then of course we can just do a

simple division to get the probability

on here. What's the probability that

we're going to roll a multiple of five

when you add them together?

but not a multiple of three. And so

we're just going to divide those two

numbers. And you can see here we get

uh.16255

or 11.62%.

And so you can really have a nice visual

that this is not really complicated math

right here on probabilities. uh it's

just how many options do you have and

how many of those are you possibly going

to be able to um come up with with the

solution you're looking for. And this

leads us to a confusion matrix. A

confusion matrix is a table which is

used to describe the performance of a

classification model on a set of test

data for which the true values are

known. And so you'll see on the left we

have the predicted and the actual and we

have a negative uh false negative

positive true positive

um and then we have false positive and

true negative. And you can think of this

as your predicted model. What does that

mean? That means if you divided your

data and you use twothird of it to

create the model, you might then test it

against an actual case for the last

third to see how well it comes out. How

many times was it uh true positive

versus uh false positive? It gave a

false positive response. And you can

imagine in medical uh situations, this

is a pretty big deal. You don't want to

give a false positive. So you might

adjust your model accordingly so you

don't have a false positive. Say with a

co virus test, it'd be better to have a

false negative and then go back and get

retested than to have 30% false

positives where then the test is pretty

much invalid. So in a use case uh like

cancer prediction, let's consider an

example where a cancer prediction model

is put to the test for its accuracy and

precision. Actual result of a person's

medical report is compared with the

prediction made by the machine learning

model. And so you can see here here's

our actual predicted uh whether they

have cancer or not. You know cancer a

big one. You don't want to have a uh

false positive. I mean a false negative.

In other words, you don't want to have

it tell you that you don't have cancer

when you do. So that would be something

you'd really be looking for in this

particular domain. You don't want a

false negative. Uh and this is again,

you know, you've created a model, you

have hundreds of people or thousands of

pieces of data that come in. There's a

real famous case study where they have

the imagery and all the measurements

they take and there's about 36 different

measurements they take. And then if you

run the a basic model, you want to know

just how accurate it is. How many um

negative results do you have that are

either telling people they have cancer

that don't or telling people that don't

have cancer that they do? And then we

can take these numbers and we can feed

them into our accuracy, our precision,

and our recall. Uh so accuracy,

precision, and recall, accuracy metric

to measure how accurately the results

are predicted. And this is your um total

um true where you got the right results.

you add them together, the true

positive, the true negative over all the

results. So what percentage of them were

accurate versus what were wrong. We talk

about precision is a metric to measure

how many of the correctly predicted

cases are actually turned out to be

positive. Uh so we have a precision on

true positive. Again, if you're talking

about like uh COVID testing with the

viruses, uh you really want this to be a

a high number. you want this true um

that to be the center point where you

might have the opposite if you're

dealing with cancer where you want no

false negatives. Uh so this is your

metric on here. Precision is your test

positive uh true positive plus uh false

positive. And then your recall how many

of the actual positive cases we were

able to predict quickly with our model.

Uh so test positive is the test positive

plus the false negative on there. And

we'll want to go ahead and do a demo on

the naive bay classifier. Before I get

too far into uh naive baze classifier

because we're going to pull it from the

sklearn or the scikit. Um let's go ahead

kind of an interesting page here for

classifiers. When you go into the

sklearn kit, there's a lot of ways to do

classification. I'll just zoom up in

here so you can see some of the titles.

Uh there's everything from the nearest

neighbor linear

uh but we're going to be focusing on the

naive bays over here. And this is just

um a sample data set that they put

together. And you can see how some of

these have a very different output. The

naive bay remember is set up as probably

the most simplified uh calculator or um

set of predictions out there. And so

what we've been talking about with the

true false and stuff like that where

there's a uh

an belief that there is a independent

assumption between the features where

the features are very assumed to have

some kind of connection uh then we can

go ahead and use that for the

prediction. And so that's what we're

using as a naive bay classifier versus

many of the other classifiers that are

out there.

For this we're going to use uh the

social network ads. It's a little data

set on here and let me go and just open

that up the file. Uh here we go. It has

user ID, gender, age, estimated salary,

uh purchased. And so we have you can see

the user ID, male 19, uh estimated

salary 19,000 and purchased zero. Uh so

it's either going to make a purchase or

not. So look at that last one. 01. We

should be thinking of binomials. we

should be thinking of simple naive base

classifier kind of setup.

So if we close this out, we're going to

go ahead and import our numpy as np.

We're nice to have a a good visual of

our data. So we'll put in our mattplot

library. Here's our pandas, our data

frame.

Uh and then we're going to go ahead and

import the data set. And the data set's

going to be we're going to read it from

the social network ads.csv. Then we're

going to print the head just so you can

see it again uh even though I showed you

it in the file. And X equals the data

set I location uh two three values and Y

is going to be the four uh column 4. Let

me just run this so it's a little easier

to go over that. Um you can see right

here we're going to be looking at uh 012

is age and estimated salary. So 2 three

and that's what I location just means um

that we're looking at the number versus

a regular location. Uh regular location

you'd actually say age and estimated

salary.

And then column four is did they make a

purchase? They purchased something. Uh

so those are the three columns we're

going to be looking at when we do this.

And we've gone ahead and imported these

and imported the data. So now our data

set is all set with this information in

it.

And we'll need to go ahead and split the

data up. Uh so we need our from the

sklearn model selection we can import

train test split. Uh this does a nice

job. We can set the random state so it

randomly picks the data. And we're just

going to take uh 25% of it is going to

go into the test our x test and our y

test and the 75% will go to x train and

y train. That way once we create our

model, we can then have data to see just

how accurate or how well it has

performed with our um prediction.

The next step in pre-processing our data

is to go ahead and do feature scaling.

Now, a lot of this is start to look

familiar. If you've done a number of the

other modules and setup, you should

start noticing that we bring in our

data. We take a look at what we're

working with. uh we go ahead and split

it up into training and testing. Uh in

this case, we're going to go ahead and

scale it. Scale it means we're putting

it between a value of minus1 and one uh

or someplace in that middle ground

there. This way, if you have any huge

set, you don't have this huge um setup.

If we go back up to here where salary uh

salary is 20,000 versus age 35, well,

there's a good chance with a lot of the

back-end math that 20,000 will skew the

results and the estimated salary will

have a higher impact than the age

instead of balancing them out and

letting the calculations weigh them

properly.

And finally, we get to actually create

our naive bay model.

Um, and then we're going to go ahead and

import the Gazian naive bays.

And the Gazian is is uh the most basic

one. That's what we're looking at now.

It turns out though, if you go to the SK

um learn kit, uh they have a number of

different ones you can pull in there.

There's a um Bernoli. I I've never used

that one. Categorical

um compliment. And here's our Gazian. Uh

so there's a number of different options

you can look at. Gazian when you come to

the naive bays is the most commonly

used. Uh so we're talking about the

naive bays that's usually what people

are talking about when they when they're

pulling this in. And one of the nice

things about the gazian if you go to

their website um to sklearn the naive

bay gazian there's a lot of cool

features. One of them is you can do

partial fit on here. Um that means if

you have a huge amount of data, you

don't have to process it all at on you

once. You can batch it into the Gausian

uh NB model. And there's many other

different things you can do with it as

far as fitting the data and how you um

manipulate it. We're just doing the

basics. So we're going to go ahead and

create our classifier. We're going to

equal the Gausian NB.

And then we're going to do a fit. We're

going to fit our training data and our

training solution. So, X-Rain, Y train,

and we'll go ahead and run this. Uh,

it's going to tell us that it it ran the

code right there.

And now we have our trained classifier

model. So, the next step is we need to

go ahead and run a prediction. We're

going to do our Y predict equals the

classifier.predict

X test. So, here we fit the data and now

we're going to go ahead and predict.

And now we get to our confusion matrix.

Uh so from the sklearn matrix metrics

you can import your confusion matrix

just as saves you from doing all the

simple math. It does it all for you. And

then we'll go ahead and create our

confusion metrics with the y test and

the y predict. So we have our actual and

we have our predicted value.

And you can see from here this is the

chart we looked at. Here's predicted.

So, true positive, false positive, false

negative, true negative.

And if we go ahead and run this, there

we have it. 653725.

And in this particular uh prediction, we

had 65 uh or predicted the truth as far

as a a purchase. They're going to make a

purchase, and we guessed three wrong.

And then we had 25 we predicted would

not purchase, and seven of them did. So,

there's our our confusion matrix.

At this point, if you were uh with your

shareholders or a board meeting, um you

would start to hear some snoozing if

they were looking at the numbers and you

say, "Hey, here's my confusion mat uh

matrix." So, let's go ahead and

visualize the results.

We're going to pull from the map plot

library colors import listed color map.

Um, and this is actually my machine's

going to throw an error because this is

being um because of the way the setup

is. I have a newer version on here than

when they put together the demo. And we

need our um X set and our Y set, which

is our X train and Y train. And then

we'll create our X1, X2. And we'll put

that into a grid. Uh, and we set our X

set minimum stop and our X set max stop.

And if you come all the way over here,

we're going to step 0. 001. This is

going to give us a nice line, uh, is

what that's doing. And then we're going

to plot the contour, uh, plot the x

limit, plot the y limit, and put the

scatter plot in there. And let's go

ahead and run this. Uh, to be honest,

when I'm doing these graphs, there's so

many different ways to do that. There's

so many different ways to put this code

together to show you what we're doing.

it's uh a lot easier to pull up the

graph and then go back up and explain

it. So the first thing we want to note

here when we're looking at the data

is this is the training set.

And so we have those who didn't make a

purchase. We've drawn a nice area for

that that's defined by the naive bay

setup. And then we have those who did

make a purchase, the green. And you can

see that some of the green dots fall

into the red area and some of the red

dots fall into the green. So even our

training set isn't going to be 100%. Uh

we couldn't do that. And so we're

looking at our different data coming

down. Uh we can kind of arrange our x1

x2 so we have a nice plot going on. And

we're going to create the um contour.

That's that nice line that's drawn down

the middle on here with the red green.

Um that's what that's what this is doing

right here with the reshape and notice

that we had to uh do the t if you

remember from numpy um if you did the

numpy module um you end up with pairs

you know x uh x1 x2 x1 x2 next row and

so forth you have to flip it so it's all

one row you have all your x1's and all

your x2s. Um so this what we're kind of

looking for right here on this setup.

Uh, and then the scatter plot is of

course um your scattered data across

there. We're just going through all the

points that puts these nice little dots

onto our setup on here. And we have our

estimated salary and our H. And then of

course the dots are did they make a

purchase or not. And just a quick note,

this is kind of funny. You can see up

here where it says X set Y set equals uh

X train Y train, which seems kind of a

little weird to do. Um, this is because

this is probably originally a

definition. Uh, so it's its own module

that could be called over and over

again. And which is really a good way to

do it because the next thing we're going

to want to do is do the exact same

thing, but we're going to visualize the

test set results. Uh, that way we can

see what happened with our test group,

our 25%.

And you can see down here we have um the

test set. Uh, and it, if you look at the

two graphs next to each other, this one

obviously has um 75% of the data, so

it's going to show a lot more. This is

only 25% of the data. You can see that

there's a number that are kind of on the

edge as to whether they could guess by

age and income they're going to make a

purchase or not. U, but that said, it

still is pretty clear. It's pretty good

as far as how much the estimate is and

how good it does.

Now, graphs are really effective for

showing people what's going on, but you

also need to have the numbers. And so,

we're going to do from sklearn, we're

going to import metrics, and then we're

going to print our metrics

classification port from the Y test and

the Y predict.

And you can see here we have precision

uh precision of zeros is 90. There's our

recall

96. We have an F1 score and a support.

And we have our precision, the recall on

getting it right. Uh, and then we can do

our accuracy, the macro average, and the

weighted average. Uh, so you can see it

pulls in pretty good as far as um how

accurate it is. You could say it's going

to be about 90% is going to guess

correctly um that it that they're not

going to purchase. And we had an 89%

chance that they are going to purchase.

Um, and then the other numbers as you

get down have a little bit different

meaning, but it's pretty straightforward

on here. Here's our accuracy, and here's

our micro average, and the weighted

average, and everything else you might

need. And if you forgot the exact

definition of accuracy, it is the true

positive, true negative over all of the

different setups. Precision is your true

positive over all positives, true and

false. And recall is a true positive

over true positive plus false negative.

And we can just real quick flip back

there so you can see those numbers on

here. Uh here's our precision, here's

our recall, and here's our accuracy on

this.

>> Welcome to this exciting journey into

the world of statistics for data

science. Have you ever wondered how data

transforms from raw numbers into

powerful insights that drive decisions?

Well, statistics is the magic behind it

all. Today, we will uncover how

statistical methods help us summarize

data, model uncertaintity, test

hypothesis, and find relationships that

can predict the future. So, buckle up.

We are about to turn numbers into

knowledge. Without any further ado,

let's get started. Now, to start off,

here's a key question. What are

statistics in data science? Now,

statistics is the science of collecting,

analyzing and interpreting data. By

applying statistical methods, we can

uncover patterns in the data and make

informed decisions. Now, as we continue,

you notice how these essential concepts

will form the backbone of many data

science practices. Let's explore the key

functions of statistics in data science.

First, statistic helps summarize data

using measures like mean, median and

variance. Next, it models uncertaintity

with probability and distributions. So,

we can better understand risk and

variability in our data. It also tests

hypothesis such as when we use AB

testing to compare different outcomes.

Statistics finds relationship through

methods like regression and correlation

revealing how variables impact each

other. And finally, all these tools

enable datadriven decision-m turning raw

numbers into actionable insights. Now,

let's talk about why does statistics

matter in data science. Statistics form

the backbone of data science, providing

the mathematical framework needed to

make sense of data and draw reliable

conclusions. Without statistics, it

would be impossible to turn raw data

into meaningful insights or make

confident evidence-based decisions. Now

let's have a look at the main branches

of statistics. Descriptive statistics

and inferential statistic. So first

let's talk about the definition. Then we

have got methods and measures. Now in

the case of descriptive statistics, now

let's have a look at the main branches

of statistics. So basically there are

two core branches. Descriptive

statistics and inferial statistics.

Descriptive statistics summarizes and

describes data using measures like mean,

median, mode, range and standard

deviation. Its purpose is to organize

and present data typically with charts,

graph or summary tables. The scope of

descriptive statistics is limited to the

sample data itself. Now on the other

hand, inferential statistics makes

inferences about populations based on

samples. It uses methods like hypothesis

testing, confidence intervals and

regression. The purpose here is to draw

conclusion and make predictions with

common examples including AB testing and

survey analysis. The scope of inferial

statistics extends beyond the sample to

the larger population. Understanding

both these branches is essential for

analyzing and interpreting data in any

data science project. Now let's take a

closer look at the descriptive

statistics starting with measures of

central tendency. The first measure here

is mean which is the average of all the

values simply calculated as the sum of

all the data points divided by the total

count. Next is the median which

represents the middle value when the

data is sorted from lowest to highest.

This is especially useful when dealing

with skewed distributions. And finally,

the mode is the most frequently

occurring value in the data set, helping

us identify common patterns or repeated

outcomes. Let's continue our deep dive

into descriptive statistics by looking

at the measures of variability. First,

we've got the range. This is simply the

difference between the maximum and

minimum value in a data set showing us

the spread of our data which measures

the average of the squared differences

from the mean. This tells us how much

the values in our data set differ from

the average. Closely related is the

standard deviation which is the square

root of the variance. It gives us a more

intuitive sense of how much the values

typically deviate from the mean. And

finally, we've got the interquartile

range or we say IQR. This shows us the

range of the middle 50% of our data,

helping us understand how data is

distributed across the center and avoid

the effect of outliers. Understanding

these four measures allows us to

summarize not just the center of our

data, but how spread out and varied our

data set is. Now let's look at some

practical applications of descriptive

statistics. One major use is the data

exploration and summarization where we

quickly get an overview and basic

understanding of complex data sets

helping track performance and detect

problems early in fields like

manufacturing or operations. And

finally, they are central to business

reporting and dashboards where concise

summaries are essential for managers to

review trends and make datadriven

decisions. So in short, descriptive

statistics help transform raw data into

clear actionable information across many

business, scientific and operational

context. Let's explore the first type of

data in statistics, qualitative or

categorical data. This type of data

includes descriptive information that

cannot be measured numerically such as

categories or labels and yes or no

responses. These are all about qualities

or characteristics rather than

quantities. Qualitative data can be

further divided into two types. Nominal

where the categories have no specific

order and ordinal where the categories

do have an order or ranking. Recognizing

and classifying qualitative data is very

important as it affects how information

is analyzed and interpreted in

statistics. Now let's discuss the second

main type of data in statistics which is

quantitative or numerical data. This

type of data includes information that

can be measured and expressed with

numbers making it ideal for mathematical

analysis. Quantitative data is further

divided into two main categories. First,

there is discrete data. These are

countable values with specific fixed

points such as the number of students,

cars sold or website clicks. Second,

there is continuous data which includes

infinite possible values within a given

range. Examples of continuous data

include height, weight, temperature or

time. Understanding the distinction

between discrete and continuous data is

very important as it determines which

statistical methods and visualizations

will be most appropriate. Now let's dive

into the fundamentals of probability.

Probability measures the likelihood of

an event occurring and it's always

expressed as a value between 0 and 1.

Here are some key concepts. If the

probability or P equals to zero, that

means the event will never occur. If P

is equals to 1, the event will always

occur. And if P is equals to 0.5, the

event has an equal chance of occurring

or not occurring. It's truly a 50/50%

scenario. Now, understanding these basic

principle help us quantify uncertaintity

and make informed predictions about

future outcomes. Now let's look at the

different types of probability. First

there's classical probability. This is

based on equally likely outcomes such as

flipping a fair coin or rolling a

balanced die. Next empirical probability

which relies on observed frequency. It's

calculated from actual data such as the

proportion of rainy days over the past

month. Let's say for example the

probability of it's raining today given

that it's cloudy. Understanding these

three type help us choose the right

approach for different situations.

Whether we are predicting outcomes,

analyzing data or making decisions under

uncertaintity. Now probability has a

wide range of powerful applications in

data science. First of all, it is used

in predictive modeling and machine

learning where algorithms estimate

future outcomes based on existing data.

Probability also plays a key role in

risk management and decision making

helping businesses and researchers

evaluate the likelihood of different

scenarios and plan accordingly.

Conditional probability calculating the

chance of one event given that the

another has occurred. This is crucial in

fields like healthcare, fraud detection

and marketing analytics.

And finally, probability is foundational

in AB testing and experimental design,

allowing us to measure the effectiveness

of new strategies or products. These

applications show how probability

enables smarter evidence-driven progress

in modern data science. Let's look at

one of the most common probability

distributions, the normal distribution.

This distribution is famous for its

bell-shaped symmetric curve, which shows

that most values cluster around the mean

with fewer and fewer values appearing as

you move away from the center. Now many

natural phenomena like heights, test

scores and measurement errors tend to

follow this pattern making the normal

distribution a key concept in

statistics. It's characterized by two

main parameters. The mean which

determines the center of the curve and

the standard deviation which controls

its spread.

Recognizing the distribution help

analysts make predictions, calculate

probabilities, and apply statistical

techniques to real world data. Next,

let's explore the binomial distribution,

which is another common probability

distribution. The binomial distribution

is discrete and is used for situations

with binary outcomes like success or

failure. It is based on a fixed number

of trials where each trial has a

constant probability of success such as

flipping a coin a certain number of

times or tracking pass fail rates.

Examples of binomial experiment includes

coin flips and measuring how many

students pass or fail a test. This

distribution help us model and analyze

outcomes when only two possibilities

exist in each trial. Now let's focus on

the poison distribution. Another key

type of probability distribution. The

poion distribution is a discrete

distribution specifically used to model

rare events. It's particularly helpful

for modeling events that occur

independently over a fixed interval of

time or space. For example, it predicts

how many times an event like a customer

arriving on a website, receiving a visit

might happen in a certain period.

Typical examples include customer

arrivals at a store, defect rates, and

manufacturing or counts of website

visits over a set period. This

distribution is great tool for

understanding and predicting random

independent events that don't happen

very often, but are important to track.

It's a branch that allows us to make

inferences, predictions, or

generalizations about a larger

population using sample data. Here are

some key concepts. First is the

population versus sample. The population

represents the entire group we want to

know about while the sample is the

subset we actually collect data from.

Next concept is sampling distribution

which refers to the distribution of a

statistic across multiple samples from

the same population. Standard error

measures how much the sample statistic

is expected to vary due to random

sampling. And finally, margin of error

tells us how much we can expect our

estimates to differ from the true

population value. Together these concept

form the foundation for drawing reliable

insights from sample data in inferential

statistics.

Let's review some of the main techniques

used in inferial statistic. First, we've

got hypothesis testing. This method

allows us to test claims or ideas about

population parameters based on sample

data helping us determine if observed

results are statistically significant

and reliability of our estimate. Next

are confidence intervals. These provide

a range of likely values for a

population parameter giving us a sense

of possible variation reliability of our

estimate. And lastly, regression

analysis is used to model and analyze

the relationships between variables,

allowing us to make predictions, uncover

trends, and understand how changes in

one factor might affect another. Now,

let's walk through the hypothesis

testing process. The first step is to

formulate hypothesis. Start with null

hypothesis represented as Hnot, which

states that there is no effect or

difference. Then there's the alternative

hypothesis represented as H1 which

suggests that an effect or difference

does exist. Now after setting up the

hypothesis the next step is to choose a

significance level noted by alpha.

Common choices for significance levels

include 0.05 5% 01 1% 010 which is 10%.

The significance level is important

because it controls the likelihood of

making a type one error also known as

false positive. And now each step in the

process is critical for ensuring that

statistical results are both meaningful

and reliable. The next step is to

collect and analyze sample data. Ensure

the data is represented by choosing a

good sample. Then calculate the

appropriate and test statistic for your

hypothesis test. Once that's done, it's

time to make a decision. Compare the p

value to the chosen significance level

alpha. Now, if the p value is less than

or equals to alpha, you reject the null

hypothesis. If the p value is greater

than alpha, you fail to reject the null

hypothesis. And finally, interpret your

results. Draw your conclusions with

respect to the context and problem at

hand. Always keeping the bigger picture

in mind. Following these step helps

ensure your hypothesis test is robust,

clear and meaningful. Let's review the

common types of hypothesis test used in

statistic. The one sample test compares

the mean of a sample to a known value

and the one sample zed test is used when

the population standard deviation is

known. Next, we have two sample test.

The independent samples t test compares

to the means of two different groups

while paired samples t test compares

before and after measurements for the

same subject. For categorical data test,

the shear test checks for the

independence or goodness of fit. And

fiser's exact test is useful for small

sample sizes. And lastly, non-parametric

tests such as the man Whitney U test and

Wil Coxson signed rank test serve as

alternatives to the t test when data

doesn't meet certain parametric

assumptions. Choosing the right

hypothesis test depends on your data

type and the specific question you want

to answer. Let's explore the central

limit theorem or CLT of foundation for

inferial statistics. The central limit

theorem states that as the sample size

increases, the distribution of sample

means approaches a normal distribution

even if the original data is a normally

distributed. This sample works for any

population distribution making it

incredibly powerful. Now for good

results, the sample size should

typically be 30 or more. What's

interesting is the sample mean

distribution which will have the same

mean as the population. The standard

error which measures variability by the

sample mean equals sigma / the square

roo of n. And this gets smaller as

the sample size grows. And finally the

formula shown here which is z is equ= to

x -

sigma divided by sigma over the square

root of n. Let's standardize and compare

sample means. The CLT makes most

parametric statistics possible and is

the backbone for many statistical test.

Let's review the main types of

regression analysis which are used to

model and understand relationships

between variables. First up is linear

regression. This technique analyzes the

relationship between a continuous

variable and another variable resulting

in a straight line. It's commonly used

in predicting sales or prices. Next is

logistic regression. Unlike linear

regression, this method is used for

binary outcomes, helping estimate

probabilities such as whether an email

is spam or a patient has a disease.

Moving to multiple regression, this

allows us to account for the effect of

several variables at once, modeling more

complex relationships like determining

house prices. Lastly, polomial

regression which is used for nonlinear

relationships. The resulting curve

rather than a straight line lets us

capture growth trends and other patterns

that aren't linear. Understanding these

types of regression help analysts choose

the right model for the data and

business's problem at hand. Let's

clarify the important differences

between correlation and causation.

Correlation is statistical measure of

how two variables move together.

Correlation values range from minus 10

to + one. But remember correlation does

not imply causation. On the other hand,

causation means that one variable

actually causes changes in another.

Establishing causation requires

controlled experimentation and is much

stronger relationship than simple

correlation. Always be careful when

interpreting results. Just because two

variables move together doesn't mean one

cause the other. Let's look at some

common problems with interpreting

correlation and causation. First is the

third variable problem which happens

when a hidden variable affects both

variables in a question leading to a

misleading connection. There's also this

directionality problem where it's

unclear which variable is causing others

to change but both are actually caused

by a third variable which is the hot

weather. This demonstrate that

correlation does not mean one variable

causes the other. So always look for

hidden factors before assuming

causation. Let's talk about statistical

errors specifically type one and type

two errors. Type one error also called

as false positive occurs when we reject

the true null hypothesis. In other

words, we wrongly conclude there's an

effect when there's actually isn't. The

probability of making this error is

equal to the significance level alpha.

For example, concluding a drug works

when it actually doesn't. On the other

hand, type two error is false negative

means failing to reject a false null

hypothesis. That's when we miss a real

effect or difference. The probability of

type two error is beta. For example,

missing the real effect of a drug and

seeing it doesn't work when it actually

does. Understanding these errors is key

for designing good experiments and

interpreting statistical results

properly. Let's look at three common

sampling methods using statistics. The

first one is random sampling. Here every

individual in the population has a equal

chance of being selected where every nth

individual is chosen. Random sampling is

crucial for ensuring representative

sample. Next is stratified sampling. The

population is divided onto homogeneous

subgroups or strata and then a random

sample is drawn from each group. This

approach makes sure all the groups are

represented in the sample. And finally,

cluster sampling divides the population

into clusters, often based on geography.

Entire clusters are then randomly

selected. It's cost effective and useful

for large spread out populations.

Choosing the right method ensures the

data truly represents the whole

population and strengthens the study's

conclusions. So guys, let's wrap up with

some real world applications of

statistics. In business analytics,

statistics are used for AB testing to

optimize websites, customer segmentation

and targeting, sales forecasting and

demand planning and quality control and

also process improvement. These

techniques help businesses make smarter

datadriven decisions every day.

Statistics play a vital role in

healthcare and medicine as well. They

are key for analyzing clinical trial

results, conducting epidemological

studies, evaluating treatment

effectiveness, and identifying risk

factors. By using these approaches,

healthcare researchers and practitioners

can improve patient outcomes in public

health. From businesses to medicine,

statistics transform raw information

into actionable insights that create

real impact. Statistics has a huge

impact in technology, data science and

finance and power recommener systems

that personalize what users see. In

finance, statistic help with risk

assessment and management, optimizing

investment portfolios, determining

credit scores, and supporting market

research and analysis. Now, these

applications show how statistical

techniques help make smarter decisions

and solve complex challenges across high

impact industry. Are you one of the many

who dreams of becoming a data scientist?

Keep watching this video if you're

passionate about data science because we

will tell you how does it really work

under the hood. Emma is a data

scientist. Let's see how a day in her

life goes while she's working on a data

science project. Well, it is very

important to understand the business

problem first. In our meeting with the

clients, Emma asks relevant questions,

understands and defines objectives for

the problem that needs to be tackled.

She's a curious soul who asks a lot of

wise, one of the many traits of a good

data scientist. Now, she ges up for data

acquisition to gather and scrape data

from multiple sources like web servers,

logs, databases, APIs, and online

repositories. Oh, it seems like finding

the right data takes both time and

effort. After the data is gathered comes

data preparation. This step involves

data cleaning and data transformation.

Data cleaning is the most time consuming

process as it involves handling many

complex scenarios. Here Emma deals with

inconsistent data types, misspelled

attributes, missing values, duplicate

values and whatnot. Then in data

transformation she modifies the data

based on defined mapping rules. In a

project ETL tools like talent and

Informatica are used to perform complex

transformations that helps the team to

understand the data structure better.

Then understanding what you actually can

do with your data is very crucial. For

that Emma does exploratory data analysis

with the help of EDA. She defines and

refineses the selection of feature

variables that will be used in the model

development. But what if Emma skips this

step? She might end up choosing the

wrong variables which will produce an

inaccurate model. Thus, exploratory data

analysis becomes the most important

step. Now, she proceeds to the core

activity of a data science project which

is data modeling. She repetitively

applies diverse machine learning

techniques like KN&N, decision tree,

knives based to the data to identify the

model that best fits the business

requirements. She trains the models on

the training data set and tests them to

select the best performing model. Emma

prefers Python for modeling the data.

However, it can also be done using R and

SAS. Well, the trickiest part is not yet

over. Visualization and communication.

Emma meets the clients again to

communicate the business findings in a

simple and effective manner to convince

the stakeholders. She uses tools like

Tableau, PowerBI and ClickView that can

help her in creating powerful reports

and dashboards. And then finally, she

deploys and maintains the model. She

tests the selected model in a

pre-production environment before

deploying it in the production

environment which is the best practice.

Right? After successfully deploying it,

she uses reports and dashboards to get

realtime analytics. Further, she also

monitors and maintains the project's

performance. Well, that's how Emma

completes the data science project. We

have seen the daily routine of a data

scientist is a whole lot of fun, has a

lot of interesting aspects and comes

with its own share of challenges. Now,

let's see how data science is changing

the world. Data science techniques along

with genomic data provides a deeper

understanding of genetic issues and

reaction to particular drugs and

diseases. Logistic companies like DHL,

FedEx have discovered the best routes to

ship, the best suited time to deliver,

the best mode of transport to choose,

thus leading to cost efficiency. With

data science, it is possible to not only

predict employee attrition, but to also

understand the key variables that

influence employee turnover. Also, the

airline companies can now easily predict

flight delay and notify the passengers

beforehand to enhance their travel

experience. Well, if you're wondering,

there are various roles offered to a

data scientist like data analyst,

machine learning engineer, deep learning

engineer, data engineer, and of course,

data scientist. The median base salaries

of a data scientist can range from

$95,000 to $165,000.

So that was about the data science. Are

you ready to be a data scientist? If

yes, then start today. The world of

data. Picture this. You're shopping

online and suddenly you see a product

that feels like it was made just for

you. How did they know? It's not by

chance. It's data science. Data science

help businesses understand what you

like, predict what you'll need next, and

improve the way we shop and use

technology. And here's the best part.

Data science isn't just about watching

Netflix. It's one of the fastest growing

careers in the world right now. In fact,

the US Bureau of Labor Statistic says

that data science jobs are expected to

grow 36% by 2033, way faster than most

of the other jobs. Companies everywhere

are using data to make smarter

decisions. That means the demand for

data scientists is huge. And let's talk

about the salary. You're probably

wondering how much can I earn actually?

Well, for entry- level position, data

scientists in the US are earning around

$152,000

per year right now. And by 2025, some

can make as much as $230,000.

And in India, starting salaries range

from 50,000 rupees to 1 lakh per month.

And experienced professionals can earn

more than 5 lakh rupees per month.

That's impressive, right? But the best

part is as a data scientist, you won't

just stop here. The skills you develop

in this role like machine learning, data

visualization, and statistics are highly

transferable and crucial for moving into

AI roles. So whether it's becoming an AI

engineer or an AI specialist, the

foundation you build in data science

will help you level up and pursue

exciting hyping opportunities in the AI

field. Now if you're thinking this

sounds great but where do I even start?

Well that's exactly what the

professional certificate course in data

science from IIT Kpur and Simply Learn

is designed to do. It will get you

started and make sure you're ready for

this booming industry. In this 11 month

live online interactive program, you

will learn the skills you need to become

a data science professional. No more

theory, no more fluff. You get hands-on

projects, life classes and mentorship

from IT Kpur faculty plus real world

industry expert who will help you build

your skills. So this course comes with

exciting amazing features that makes

learning even more impactful. Eight

times high interaction in live online

classes with industry expert. Regular

live online classes conducted by

experienced professionals who bring real

world knowledge into every session.

You'll also get to master 13 plus key

skills including generative AI, prompt

engineering, charge, expendable AI,

conversional AI, NLP, and many more.

These are the skills that top companies

use every day. And you'll gain hands-on

experience with 14 plus industry tools

like Python, SQL, Tableau, Dali 2,

Midjourney, TensorFlow, and more. And by

the end of this course, you will be

ready to tackle real world challenges

using these powerful tools and

techniques. But we are not just talking

about textbook and theory. You'll work

on 25 plus real world projects giving

you hands-on experience with the tools

and skills you'll use in the industry.

For example, our first project would be

about sales analysis. You'll use Python

to analyze a clothing company fourth

quarter sales data across Australian

state, helping the company make informed

decisions. The next project would be

about employee performance analysis

where you learn how to build machine

learning models to understand the

factors influencing employees turnover.

Our third project would be about

e-commerce which will help Amazon

improve its recommendation engine to

offer better recommendation to

customers. Upon successful completion,

you'll receive a program certificate

directly issued by the ENI city academy

IT Kpool within 45 days of completing

your cohort. This prestigious

certificate will help you boost your

resume and show potential employees that

you have mastered the skills needed to

succeed. You'll also benefit from master

classes delivered by distinguished IT

Kur faculty who bring their deep

expertise in this course. Plus, you will

be exposed to trending tools like

chargeb2 geni and prompt engineering.

Now, if you're wondering who will be

teaching all of this, then the IT

carpool faculty is here for you. These

experts have been in this field for

years and have worked with the top

companies. You'll also get master

classes from them and they will guide

you through the learning process. You're

not just learning from a textbook. You

are learning from the people who have

been there and done that. And the best

part is once you have learned the

skills, simply learns career assistance

team will help you take to the next step

which will help you to build a killer

resume, show you how to stand out top

recruiters and even give you access to

mock interviews. Plus, you will also

gain access to exclusive networking

events and hackathons to connect with

industry professionals. Along with

Simply Learn's job assistant, you'll

also get access to ID Kpur's career

services helping you connect with top

recruiters and land interviews with

leading tech companies. So, upon

finishing this course, you'll be ready

for top data science roles like data

scientists, machine learning engineer,

AI specialist, and business analyst. The

good news is top companies like Amazon,

EY, Fidelity Investment, Johnson and

Johnson, Borafhone, Accenture, Infosys

and Nvidia is looking for professionals

just like you. And with salaries in the

US hitting around $230,000 plus and in

India reaching up to five lakh per

month, your career outlook looks great.

So what will you be actually learning in

this course? So here's a sneak peek of

the syllabus which you'll be learning in

this course which is the foundation in

Python, SQL and mathematics, core data

science like machine learning, data

visualization, NLP, special topics like

GNI, charge GBT, prompt engineering and

you'll also work on industry projects

which we have already mentioned before.

You'll also have the option to choose

electives like data storytelling with

PowerBI and business analytics with

Excel to tailor your learning experience

and focus on the areas that interest you

the most. So don't wait, hurry up and

enroll now and find the course link in

the description box below and in the pin

comments. Have you ever wondered how

your favorite online store seems to know

exactly what you are looking for? Every

time you browse, add to cart or wish

list an item, you are leaving clues

about your style, favorite colors,

brands, and even shopping times. Data

scientists jump in, analyze these

patterns, and create a super

personalized shopping experience.

Suddenly, the store is showing you just

the right pieces at just the right time.

Almost like it's reading your mind.

That's data science. Turning your clicks

into a shopping spree crafted just for

you. Hello everyone. Welcome back to

Simply Learn's YouTube channel. If

you're already a data science enthusiast

or just got curious about this exciting

field, you're in the right place. Today

in this video, I'm diving into 10

essential steps to help you become the

next in- demand data scientist and land

that dream job. No more waiting. Let's

dive right in and get you on the path to

your future in data science. So, let's

see the 10 essential steps to become the

next data scientist in demand. Step

number one is programming languages.

Starting with Python is a beginner is a

great move because it's simple,

versatile, and widely used in data

science. Python straightforward syntax

makes it beginner friendly, helping you

grasp programming basics quickly and

dive into data science libraries like

pandas, numpy and mattplotive with ease.

Adding R to your skill set is valuable

because it excels at statistical

analysis and data visualization two

essential parts of data science. You can

be comfortable with Python and R within

a month or two. So moving on to the next

step that is version control system.

Learning a version control system like

Git is essential because it allows you

to track, manage, and collaborate and

code effectively. With Git, you can save

different versions of your work, making

it easy to backtrack if something goes

wrong or to experiment without losing

progress. This is especially useful when

working with complex data science

projects where you might try out

different models of analysis techniques.

One or two weeks for practice along with

Python and R is good to get start. Now

moving on to the third step that is data

structures and algorithms. Learning data

structures and algorithms is crucial for

becoming a data scientist because they

provide the foundation for efficient

data handling and problem solving. Data

structures like arrays, stacks, cues and

trees help you store and organize data

in ways that make it easier and faster

to access, process and analyze.

Algorithms on the other hand give you

strategies to perform tasks like

searching, sorting and optimizing data

operations which are essential for

handling large data sets. While many

candidates struggle with the essay,

mastering it gives you an age helping

you stand out in the interviews and

shine as a skilled data scientist

capable of tackling the toughest data

problems. Spend about two months in

this, you will get in the shape for

sure. Now moving on to the step number

four that is SQL. Learning SQL is

essential for data scientists because it

enables you to access, manage, and

manipulate data directly within

databases where most real world data

resides. With SQL, you can create new

tables, alter existing ones, delete

unnecessary records, and run queries to

filter, sort, and aggregate data. These

abilities allow you to retrieve, clean,

and organize data effectively. Core

skills needed for any data science role.

It's easy and you don't have to spend

more than a month to have a deep

understanding of it. Now moving on to

the fifth step that is mathematics and

statistics. Mathematics and statistics

are essential for data science because

they form the backbone of data analysis,

model building and interpretation.

Topics like linear algebra, calculus,

probability and statistics gives data

scientists the tools to understand data

patterns, perform accurate analysis and

make datadriven decisions. Mastering

these areas enables you to build robust

models, validate results and tackle

complex problems confidently making you

a well-rounded and skilled data

scientist. Make sure you spend two

months to grasp this topics. Now moving

on to the step number six that is data

prep-processing and visualization.

Learning data prep-processing and

visualization is essential for a data

scientist because these skills make you

data accurate, insightful and easy to

understand. Python libraries like NumPy

and Panders are crucial for manipulating

and creating data, enabling you to

handle missing values, filter out noise,

and prepare data for analysis. Once the

data is ready, visualization lets you

uncover patterns and communicate results

effectively. Libraries like Mattplot tip

and Seaborn help create clear, impactful

visuals, allowing you to interpret

trends and convey insights in a way

that's easily understood by others.

Together with these tools make data

prep-processing and visualization

fundamentals for effective data science.

If you have a solid foundation on Python

and mathematics, you will get a good

understanding of data prep-processing

and visualization in a month or two. Now

moving on to the seventh step that is

machine learning fundamentals. Machine

learning fundamentals involve

understanding how algorithms enable

computers to learn from data and make

predictions on decisions without

explicit programming. The two main

categories are supervised learning and

unsupervised learning. In supervised

learning, models are trained on labelled

data to make predictions while in

unsupervised learning models find

patterns in unlabelled data. Popular

tools like TensorFlow, PyTorch help

build and train complex models

especially for deep learning. While

skyit learn is essential used for

simpler machine learning algorithms and

data prep-processing. These tools make

it easier to implement machine learning

fundamentals effectively and build

intelligent datadriven decisions.

Dedicate about three months to

understand the core of machine learning.

Now coming to the next step that is deep

learning. Deep learning is a subset of

machine learning that focuses on

algorithms inspired by the structures of

the human brain called neural networks.

Deep learning uses neural networks with

multiple layers often dozens or hundreds

to learn complex patterns from large

data sets. Specialized types like

convolutional neural networks that is

CNN's are great for image processing

while recurrent neural networks RNNs are

used for sequence data like text or time

series. Essential tools like TensorFlow,

PyTorch make building, training and

deploying deep learning models more

accessible, allowing you to create

powerful AI solutions across various

domains. I think it will take about 2

months to have a good hold on deep

learning concepts and how to implement

them. Now moving on to the ninth step

that is specializations. Once you have

grasped the deep learning, it's like

reaching a new level as a data

scientist. Just as doctors specialize in

areas in nephrology and cardiology, data

scientists often choose to specialize in

fields like natural language processing

or computer vision. Natural language

processing focuses on teaching machines

to understand and generate human

language enabling applications like

chatbot, sentiment analysis, and

language transition. It's about making

computers read, write, and even

interpret human emotions through text or

speech. Computer vision on the other

hand is all about enabling machines to

see and interpret images or videos. This

field powers innovations like facial

recognition, object detection and

autonomous driving. Now you don't need

to learn both. You can choose what

interests you the most. Now spend one to

two months diving deep into one of these

areas. Now moving on to the last but not

the least step that is big data. Big

data refers to extremely large volumes

of data generated rapidly from sources

like social media and sensors. For data

scientists, learning to handle big data

is crucial as it requires specialized

tools like Hadoop and Spark to analyze

and extract insights effectively. With

companies relying on datadriven

decisions, big data skills make you a

highly in- demand professional in the

field. Focus for about 2 months and you

will be able to spot trends and patterns

from data sets very easily. Once you're

ready, it's time to build a killer

resume packed with projects that

showcase your new skills. Start applying

to jobs on platforms like Noy and Indate

and supercharge your LinkedIn. Connect

with data scientists. See what skills

they are mastering and learn from their

journeys as well. Keep sharpening your

own skills and when the time comes, you

will be ready to crush those interviews

and land your dream data scientist role

in 2025.

>> Yeah. step by step we will go through

all of this and uh we'll make sure that

we learn everything and we bring

everything together towards the end

right without further ado let me just

straight away deep dive to business

right to learn data science

right

and with this data science there is also

something which is prefixed which is

applied data science

and suffix for this is with Python

right apply data science with Python

right so there are there are two key

concepts which are going to be a part of

this course the first one is the

knowledge about data science that what

data science is and then because we are

doing an applied course right we are

doing an applied course I will try to

tie up these concepts which we will

understand in data science with a tool

right which is Python for you right we

already know about uh 60 65% of Python

right which is the fundamental Python

and now we will be moving to the next

step to advanced Python

right and using Python right leveraging

Python we will be solving a lot of

problems of data science right using

this.

Okay. So, the first few sessions, right?

The first few sessions will be about

making you a breast with Python. What

Python is, right? What how and what

packages do we have? How do they work in

reality, right? And all those things.

And then we will be coupling it up with

data science concepts. And then finally

towards the end of the session in the

last few classes, we will be doing uh we

will be taking a real data set. And on

that data set we will be applying all

these concepts right to understand the

data better and we will be drawing

inferences from that to convert that

into information to take actionable

insights or using that actionable

insights taking a better decision.

Right? We'll do all of that in in the

actual way. Okay. So now guys if you

understand this then the next point of

contention is data sets

right one of the most famous keywords on

the planet right now right one of the

most famous keywords in the planet right

now do you think that these two things

okay let me put it different way what do

you think that can be the possible

explanation about this term data science

you know data You know science what do

you think is going to follow in these

sessions? What is data science to you as

per these two words? Okay. So this is

people made up of two words right data

and science right. So what we are trying

to do is we are

trying to understand data

right? We are trying to understand data

right and then do something to it right

understanding its science understanding

the uh nature the behavior of this data

and converting it in something called as

information

right do we know difference between data

and information data is something which

is completely raw okay it is completely

raw it has no meaning

right it has no meaning isn't it for

example

I give you these stats of some player

like suppose Sid Dhoni I give you stats

of Mahindra Singh Dhoni right that what

what what were his scores uh what is his

name what is his age and you know all

those things now everything is there

right but we don't know what to do about

it right do you think the score of Dhoni

has any context people it has any

context no right but when I deep down

but I when I go and deep dive about it

right what is the first thing you find

out of scores what is the first thing

you find out of score scores you try to

find the average of score isn't it that

in last 10 innings

right in last 10 innings before this

also you need something which is called

as a problem statement isn't it now for

example the problem statement is select

Selectors want to understand selectors

wants to understand that whether

Mahindra Singh Dhoni should be picked

up. So there's a problem now right

selectors want to see that whether Dhoni

is fit for the next tournament or not.

So what we will do we will now try to

take the mean of the scores right for

last 10 innings. And if this score is

suppose X, we will try to compare this

with Y. What is Y? Y is a reference,

right? Y is a reference that we want to

compare it against. Now when you are

doing this comparisons, when you are

applying these techniques to this score,

this is now slowly becoming information,

right? And at the end of the day once

you have the strike rate once you have

the mean score of Dhoni once you have

his age once you have his fitness score

all those things will now help you to

take this particular decision because

now what you have is called as

information right because this has

context

right this has meaning

and this is usually processed

Right? This is usually processed. Right?

This is usually processed. Now what did

we do? Now what did we do here? If you

will go and read about data science,

data science says,

data science says

it is the art of collecting,

right? cleaning,

analyzing,

modeling,

improving,

right? And visualizing,

right? Visualizing

the day, right? If a person is adept in

doing all these things, this person

people is cumulatively called a data

scientist. Right?

That person is called a data scientist.

Right? So before going to the definition

of data scientist, now I will give you

some more examples, right? I'll give you

some more examples. Data science people

as I said is a combination of these

things, right? You have to collect the

data,

right? You have to collect the data,

right? Right. And this has a lot of

things. Data can be connected from two

types in two types. One is primary

and the second one is secondary.

Right? What is the primary way of

collecting data? From your IoT devices,

right? From sensors,

from your inbuilt machines,

right? Then from your surveys

which you float, right? Questionnaires,

right? All these things are primary

ways. What is the secondary way of

collecting data?

Purchasing data,

right? Using internet data

because you have not generated it. You

are just using someone else's data.

Right? Something like uh transfer

learning.

What is transfer learning?

Transfer learning is a technique where

suppose I am bank A and you are bank B

right so bank A has created some model

right trained on their data now you are

going to use the exact same model right

you're going to use the exact same model

right maybe you're not seeing the data

but you are just using the property of

data like mean median mode and a lot of

modeling things which will come we will

learn about them and You use this model

on your particular data right so in a

way you did not have enough data to

create the model yourself but you are

now using someone else's model to run

your data on it right so this is called

as transfer learning so this kind of

collection is basically secondary data

collection so you can collect the data

right then you can perform data analysis

right you can perform data analysis

right How will you perform this data

analysis? Using complex

algorithms,

right? Using complex algorithms, right?

Some statistics,

right? You can use artificial

intelligence,

right? Artificial intelligence. You can

use machine learning.

Right.

Right. You can use all these things for

data analysis. Then you can transform

transform

the patterns

into predictions.

Right? You can transfer these patterns

into predictions, right? Which can be

used for business

decision making,

right? For business decision making.

Then you can validate the results,

right? And present the results,

right? So this is like a complete life

cycle of a data scientist, right? So

before going further, let me give you

what combinations do you need to have to

become a data scientist. The first one

is

domain knowledge,

right? So what is domain knowledge?

First of all, I told you right there

will be a problem, right? You'll be

solving a problem in any project of data

science. You'll be trying to solve a

problem right and the problem will be

belonging to a particular domain even if

you're working for yourself right even

if you're an entrepreneur then also

you'll be solving a problem. So this

domain knowledge part includes things

like understanding

it's a very important diagram

understanding

client requirement

right understanding the client

requirement right

important criterians

right important

criteria knowledge

Right? For example, to give an example,

suppose we have created a machine

learning model. Okay? Understanding the

data, we have created a machine learning

model whose accuracy is 90%. Right? Is

90% a good accuracy?

Yeah, fairly decent accuracy. Yes.

Suppose you have to predict sales,

right? You are selling something.

Suppose you are selling clothes and you

want to predict what will be the sales

for the next week. When you use this

model, whatever the output model gives

you, what is going to be the accuracy of

your output using this model?

How much accuracy?

90%.

But my question is model is 90%. But my

question to you is that if 90% accuracy

is on sales data, a person like me will

be very very happy. Okay? Very very

happy. I'll be probably dancing, right?

But if you try to apply the same model

right for a medical diagnosis case, will

you be interested in getting operated in

such a hospital or an institution where

the accuracy is coming as 90%. Domain

knowledge, right? Domain knowledge. We

need to understand what are the exact

requirements. We need to understand what

are the exact expectations,

right? And we need to know how much do

we need to pivot right? So the first

thing in data science is these

accuracies and everything are subjective

right they are subjective. So for that

you need domain knowledge. Domain

knowledge part very important guys very

important these three things which I'm

going to tell. Second part is people

the game changer right the second part

is

computer science.

Now if I take you back in history okay

if I take you back in history in 1980s

or somewhere then do you okay how many

of you think that data science is a new

concept

how many of you think that data science

is a new concept I hope you all it's not

a new concept everyone knows that yeah

it has been happening for ages just like

you guys will be shocked if you already

don't know AI was coined in the year

1956

1956 6 at the University of Dharma. AI

was coined by Paul McCarthy, right? And

we saw the boom of AI in the year 2010,

right? Such a long journey. Same case

with data science because people back in

the day data science was called as data

mining. Everyone heard about it data

mining, knowledge databases. Yeah, we

need we used to mine the data. Now, what

were the problems? What were the hiccups

of data mining? The hiccups for data

mining was that we were doing everything

everything manually

right now if I give you 100 points can

you calculate the mean

or let's say if I give you two points to

multiply 2 * 3 how much time will you

take?

2 seconds.

Yep. How much time a computer will take?

2 seconds. If I give you to multiply 2

489

multiplied by 200, how much time will

you take to calculate this? Say 5

seconds. How much time computer will

take? 2 seconds. Now if I give you to

multiply 2 48 9 into 15 1 95 4386

how much time will you take to calculate

this manually? Maybe say 1 minute

80 seconds 1 minute. How much time a

computer will take? Still 2 seconds

right? still 2 seconds right and if I

give you to calculate this over 200

times you will take 200 minutes right

using parallel computing computer will

still take about 3 to 5 seconds right so

are you understanding the power of

computer do you understand this concept

in this relationship what was happening

back then what computer science did

people was it revolutionized the way

data mining was happening and That thing

now is called as that again that thing

now is called as data science in which

computer science is one of the most

important contributors. So this is just

one reason. Now manually right manually

if I give you say 1 million rows of data

right 1 million rows of data right so

how many pages will you pages will you

need to store this data suppose your

notebook is like this

right these boxes and here you are

storing the data 1 million so maybe you

can buy n number of notebooks but now do

you think it's as easy in storing

something in computer because back in

the day people we had memory issues,

isn't it? Memory constraints.

So there is something called as Murray's

law, right? Which says as the

advancement in microprocessors will

increase, the price of microprocessor

will decrease, right? So this is what is

happening right now. Back in 1980s, if I

show you right guys, right? Yeah. 2.5

kg. Exactly. Right. It was size of a

fridge hard disk but now it fits in your

palm right. So this was enabled the

storage techniques right the processing

techniques the infrastructure right

things like big data what kind of data

do you think we will be dealing with

people in data science you all know the

term very famous term the kind of data

big data right everyone knows about big

data what is big data yes a data which

is fast right it has velocity veracity

variety right so This kind of data needs

to be stored. This kind of data needs to

be processed. So which thing brought all

these things into data science? It was

given to us by computer science, right?

So computer science people included

things like database management,

right? Data validation, right? Data

infrastructure,

right? Data infrastructure,

right? Then we had languages, computer

languages

which is Python right now for us. Right?

Again do you think when you do this

thing manually right suppose you do this

thing manually how easy do you think it

will become using something like Python

or any other computer language to create

complex models. How easy it will be to

do that to create the complexity in

models right where you can capture the

nonlinear nature isn't it people isn't

it

for example for example let me tell you

this

2 4 6 8 10 dash what do you think is the

next number guys 12 if I tell you to

define this to me in okay let leave

leave

What do you think is going to be the

next number here?

What is the next number? 11.

Next number

25.

Perfect. Now guys, if I ask you to write

these numbers, right, the way you

predicted them, can you give me a

function f ofx is equal to what is the f

of x here?

It's 2x, right? It's 2x. f ofx is equal

to 2x. If I tell you to create a

function, it will be f of x is equal to

2x. What will be the function here,

guys?

f ofx

will be equal to

x + 1. Yeah. x + 1. Yeah. And here f ofx

will be equal to

xยฒ. Yeah. Now the last example, right?

Last example.

What is the next number here? I don't

want the number. I want the function. I

want this so that I can generalize.

Isn't it? How did you reach this figure?

How many of you think it's not possible

to determine this? How many

of you

think

it is not

possible

to determine this?

Yeah. How many of you think what if I

just change this question and ask you

how many of you think it is not possible

to determine this

manually?

Same response. But now I say how many of

you think it is not possible to

determine this

with computers?

Will your answer still remain no? Do you

think I cannot approximate this function

using computers?

We have something called as deep

neural networks

and they are called as universal

function approximators.

Right? So this is the problem people.

This is the problem. Right? I will show

this to you when the time comes. Right?

I will remember this example and I will

show this to you. But now what I'm

trying to tell you is the things which

seemed impossible manually was solved by

what? It was solved by computers. The

distribution of this is like this. Can

you figure it out yourself? No. Right?

We cannot. Isn't it? We cannot do that.

So this kind of approximation will be

given by what? It will be only given by

machines. Right? And this is people what

data science is all about. Right? It is

what computer science did inside data

science. Right? I hope this is clear.

I'm assuming a lot of you will be going

for interviews and everything after

these course. Right? So this will be a

very very important thing for you to

know. Right? Often it is asked why data

science is having computer science in

it. Right? The reason is this. Okay? So

this is the role of computer science

inside data science. Now people the

third thing the third circle which is

one of the most

parts is

maths

and stats right mathematics statistics

which was optimization

right optimization of your models right

design

of model. Right? Now guys, if you look

carefully, if you look carefully, in

order to approximate this, what you what

will you be playing with? You will be

playing with a lot of data. You'll be

playing with a lot of mathematical to

mathematical concepts and statistical

concepts, isn't it? How did you what do

you call this? This is math, right? This

is statistics and mathematics. finding

mean, median, mode, standard deviations,

probability, statistics, all these will

lead to this kind of result, isn't it?

So, this becomes the third wheel of this

particular uh of of this particular

diagram. And this point of intersection,

right? The sweet point of intersection

is basically data science.

Yeah. is particularly data science.

Right? So now people this point okay

this point

is basically representing data

engineering right data engineering right

data engineering is people the part of

data science which enables us to capture

the correct data right how the data will

flow how the data will be stored how the

data will be cleaned right all this is

done by home it is done by data

engineering Just because so am I right

there you'll go for interview right

after this and try to fetch yourself

jobs in this domain data science AI ML

if my understanding is correct is that

the aim yes so now guys there will be

three types of companies

or let's say to simplify let's say two

types one is small and the other one is

big right so in a small organization if

you become a part of small or

organization and you are the data

scientist there You can be involved in

all of these things, right? All of these

things possible, right? Your bosses and

your management will expect you to

construct all these flows, right? Know

computer science, you should know maths

and stats and you should have the domain

knowledge and you will be asked to do

all of this. But if you are going to

become a part of a big organization,

usually all these roles are fragmented.

All of these roles are fragmented,

right? There's a separate data

infrastructure team. There's a separate

data governance team. Now guys, when you

go on to collect the data, can you

collect any sensitive data about is it

possible ethically it's not right? And

legally also it's not right. My question

to you is who will look after this

compliance? Whose responsibility indeed

it is to look after this compliance?

Data scientist. So this is about

fragmentation. If you are part of a big

organization, this thing will be taken

by someone else, right? But if you are a

part of a small organization, you will

be know you'll be expected to do this

all by yourself. In if if you are part

of a big organization then do you think

you need to have this domain knowledge?

The answer is no. Why? Because there

will be separate set of people who are

called as what? Who are called as

business analyst. Have you heard about

this position people? Business analyst.

What is a business analyst role? It is

basically a technical translator, right?

who knows technical, who knows domain

and that person goes and talks to the

client, talks to the client in a layman

language, convert it into technical

requirement coupled with the domain

knowledge and give you the document.

This is basically a medical engineering

problem. So we need this this this this

and you need to fulfill this this this

this criteria. But again, if you're part

of a small organization, who needs to

take care of that? Who needs to make

sure that you know everything about a

domain? You yourself, right? You

yourself, right? Then people, this area

usually represents whom? This area

represents research

and analysis.

Why?

Because they we have people who have

domain knowledge and we have people who

are knowledge of math, stats. Have you

heard about a position called as actury

in the world? Acturial science. Acturies

are people who are basically dealing

with uh domains which are very very

heavily data intrinsic. Right? For

example, finance domain, right? Finance

is all about numbers. So there we go

actal science and we have math, stats,

optimization, model development, all of

those things happening there. We are

also at at this point of time people

belonging to which section research and

data analysis right and then people

there's the third intersection right

there's a third the third intersection

which is this part and this people is

called as machine learning

right machine learning why machine

learning if you can combine the power of

maths stats right and you Combine the

power of computer science, you will find

yourself to be in a position where you

can call yourself a machine learning

engineer. Right? How does a machine

learning engineer becomes a data

scientist? When they couple it up with

the domain expertise, right? So in this

course people in this course we will

teach you computer science. We will

teach you a little bit about math stats.

But what we cannot teach you is domain

knowledge. Yeah.

and making the base of what we are about

to do. Very very important to

understand. Right? If you understand

this then half the battle is won. Right?

So now to answer the question which was

posted earlier uh answer to the question

which was posted earlier. There are

different things in the world of data

science. Right? So you can pick and

choose anything or you can do everything

by yourself. If you guys are engineers

then I think you can be at the sweet

spot going forward in life if you choose

a domain for yourself. For example, you

choose to be a part of automobile

industry, you choose to be a part of

medical industry, you choose to be a

part of say retail industry, you choose

to be a part of aeronautics industry,

right? You choose to be a part of

finance industry, right? So whatever you

will choose, this thing will get

developed over time, right? This is the

most difficult out of these three. I try

to give you the example right my domain

was agricultural industry right the agri

products I have worked extensively in

agricultural industry right so again for

now in my current role this is something

which I don't have I have this expertise

I have this expertise so same will be

with you and you guys will develop this

knowledge over the time now I will try

to give you an example right elections

to make you understand how data science

can be used in one particular use case.

Okay, maybe we can extend that to a lot

of other use cases and examples, right?

Talking about election season, right?

Talking about the election season,

right? We will we will try to understand

how do we use data science because it is

very extensively used in this data

science. Right? Now, let me talk about

the first phase. Let's say this is

pre-election

phase,

right?

Right. This is pre-election phase. In

this pre-election phase, what do you

think will be the tasks with which an

agency like Election Commission of India

will be doing? The first task can be

that they will be

doing the voter

registration,

right? Voter registration and data

management, isn't it?

Yeah, it will start with that.

And what will be the things inside this?

The first thing will be data collection,

right? First thing will be data

collection. So, you will collect the

data from all the registered voters,

right? Maybe it can be their demographic

information, where they live, what is

their age, what is their gender, right?

What is their past polling behavior?

Have they turned out previously or not?

Right? All these details we can collect.

Then can we also do data cleaning?

Because I've told you, right? That there

can be a lot of redundancies. Some

person's name can appear twice, right?

Some people can be a mismatch. Suppose

for example, we have learned this in

Python. Raghav.

Raghav.

Raghav.

Radha. Right.

Right. All these are what people?

This is belonging to the same name.

Right. This is me. But for a computer,

for a computer, how many ragavves are

there? Different ones. All are

different. Right? So example like these,

right? Some people might have died. They

might not be existing anymore. Right? So

all this part will be taken care where

people in the data cleaning process.

Right? Removing the duplicates, updating

the new addresses, right? Correcting the

information about every voter, all those

things, right? And now lastly, we can

also include a flavor of data analytics,

right? What will data analytics include

in this? We can analyze the demographic

data to identify eligible voters. Isn't

it? Yeah. People till now we haven't

understood this why it is not automated.

How will you pick up all the how how

will you pick up all the details and

nuances? Suppose you're filling a form

right by mistake you have. Suppose you

are 25 years of age. Suppose you have

written 250. So does that mean I remove

this? I remove this entry of yours

because by mistake you have written your

age as 250

is age 250 possible in our current world

never right so I will have to tell the

machine that because this is a mistake

please convert it to 25 isn't it this is

called as imputation so this is mostly a

manual task not manual but you have to

understand the problem manually and then

code it on

Suppose someone has written their state

as

E D L H I right and country

as I D I N right so what is this state

referring to what is this country

referring to the humans are very smart

India right we can say this is India and

if this is India then what is this

pointing out to

telly just because you said and it's a

it's a leading question I want you to

explain this tell me is it possible to

do this automatically no right we have

to employ manual rules right we have to

tell because we who is more intelligent

humans or machines humans right machines

are just more optimized right so we know

through human intelligence that this is

pointing to Delhi and this is not EDLHI

so this is data cleaning Right. Lastly,

we have data analytics. So, do you think

people based on these data points, we

can understand that who are the eligible

voters

and maybe who are not registered yet,

maybe who have not voted in the past.

Can we do all those analytics

and reach out to those peoples and

persons? That's the first part. This

just the first part pre-election phase.

Now moving on to the second part right

moving on to the second part let's say

uh we say public

opinion

regarding the polling right the polling

which is about to happen the first thing

will be you want to collect information

about people right so can can you go out

and reach all 1.8 8 billion people in

this country that what is their likable

vote for which party is it possible

1.8 8 billion do you think it's possible

for 1 billion

do you think it's possible for 500

million do you think it's possible for

100 million no right so basically I'm

talking about what I have something

which is called as a population

right and if I have to study about this

population which is 1.8 8 billion people

which is impossible which you just said

what do I need to do should I stop my

process no right I will go and collect

something which is called a sample

right we always work in samples right

suppose someone says that a Coca-Cola

bottle does not contain 500 ml of liquid

which it claims right suppose someone

has put this allegation possible that

Coca-Cola bottles do not have 500 ml

liquid which they promise. Now there are

two ways to deal with this. Right? There

are two ways to deal with this. Either I

go and collect all the bottles of

Coca-Cola in the world. Possible

never right. So what will I do? I will

go and pick up handful of bottles.

Right? Handful of bottles. So what is

that handful of bottles? Those are

called as samples. One last thing.

Suppose someone says that because of an

industry

all the fishes of the lake are dying or

they are infected. Is it possible to go

and collect and check all the fishes in

the pond or a lake? No. Right. What will

we do? We will collect again handful of

fishes and we will test them. Right?

Again samples. Now how does the raw data

collected? Raw data as in I hope you

understand this. This is no more about

population

with this step being told. Now you're

dealing with samples. So now do you want

to ask me how is sample created? Yeah,

now I'm coming to that. Now guys, there

are a lot of second point is how to

sample right? How to sample

right? So we have sampling techniques

people. One is called as probabilistic

and one is called as nonrobabilistic.

Right? I will not go in detail right

now. I just want to tell you an overview

probabilistic is suppose uh you are

manufacturing t-shirts right you are

manufacturing t-shirts right and suppose

you created

100 lots

of

thousand t-shirts right so this is box

one box two box three box four up to up

to 100 right 100 boxes and in each box

how many t-shirts are there 1 th00and

right now suppose you are a Quality

inspector. You're a quality inspector.

Is it possible you for you to go over

all the 100 lots with all checking all

the thousand t-shirts one by one? No.

Right. What will you do? You will sample

again. You will sample. Now the most

common way of sampling these kind of

problems is probabilistic sampling. What

is probability? What is the probability

of getting heads or a tails when you

spin the when you flip the coin? equally

likely 1x2 and 1x2. What is the

probability of getting 1 2 3 4 5 6 on a

roll of a dice? 1x 6. Now what is the

probability of picking any t-shirt from

this first slot out of thousand

t-shirts?

1 by,000.

Yes. So do you think all the t-shirts

have equally probable equal probability

of being picked up without any bias? If

you decide to draw five t-shirts, right,

from each of this box, right, and

suppose say two are defective and three

are not defective, what will you do?

Will you accept the lot or reject the

lot? We have majority of t-shirts of

non-deective

or let's say we will reject the lot. We

will reject the lot. We will reject the

lot. Let's say we will reject the lot.

Okay. Though this is basically

subjective as per the company policies

but let's say we rejected. Now people my

question to you is what if this entire

batch had only two defected t-shirts but

now what will happen? The entire batch

will be rejected.

Yes. Let's say you sampled one t-shirt.

Let's let's change the use case. I say

you sampled only one t-shirt and that

t-shirt was defected. Now you will

reject the batch and that is equally

likely case. So this is called as

probabilistic sampling people and there

is no way you can go back. There is no

way you cannot say that hey sir please

uh allow this batch to pass because

there is a chance that rest of the

t-shirts are not defective. No it is not

the way that happens. It happens

randomly. So right this is called as

random sampling.

Right? random sampling. Now suppose you

are doing a cancer research, right?

You're doing a cancer research, right?

So for your cancer research people, what

kind of people will you need? People who

had had cancer in the past, isn't it? To

know more about their problem, to know

more about their medical condition. So

is it possible people that in this use

case you can go and pick up any person

from the population and ask them

questions? No. Right? That is not

possible. So now is the probability

equally likely or it has changed when

you pick the sample? It has changed. Now

there is a bias which is introduced that

you only want people who had cancer.

Right? So that kind of sampling people

is called as nonprobabilistic sampling.

Right? Non-robabilistic sampling. Clear?

Now sampling technique. Right? Now third

thing in this same scheme can be people

what? It can be the data collection

right data collection mode right that

how do you collect the data? You can

float a survey

on say internet.

You can go and stand outside a mall

or office,

isn't it? How will you how will you can

probably interview someone,

right? Interview someone, right? You can

have a group discussion.

Yeah. All these techniques.

Yes. No, maybe. Right. In the same part,

public opinion polling, right? Now,

guys, uh so this was a brief

introduction, right? And this can be

extended to any industry. Right? As of

now, you can have example in the medical

science,

right? You can have an example in

automobile,

right? You can have an example in

retail,

right? Right. Then you can have example

in manufacturing.

Right? You can have example in

education,

right? You can have example in sports,

right? IPL analysis, cricket analysis,

all these sports analysis, right? These

are the most famous domains, right? They

are the most famous domains. They are

not topics, they are domains, right? In

which data science is used extensively.

So, I'm just going to check. Guys, in

automobiles, there's a biggest example,

self-driving cars.

Yeah, autonomous driving. How do you

think that's possible? Data science

again

like Tesla. Absolutely. Tesla is level

three. We have five levels.

Level three is narrow AI. Level four is

AGI and level five is super AI. We are

going to move first right with the

technical aspect right with the

technical aspect in our data science

course right

right which is based on Python

right because that is our base language

which we have learned so far right so in

Python people we will start and cover

four of the packages

right now and as we move on to machine

learning and other uh deep learning and

everything you will explore more and

more packages. The first package we have

to cover will be numpy.

Right? I'll explain you in detail what

numpy is. Then we will cover pandas.

Then we will cover mattplot lip

and then finally we will cover something

called as cbond.

Right? We will cover something called as

cbond. So these four packages inherently

we have to cover in Python to make sure

we are able to reduce

the time

in coding right we are able to reduce

the time in coding and using these

packages immediately help us in getting

the desired results right I hope you all

remember the concept of modules

we have covered in Python do we all

remember functions and modules. You can

use Jupyter notebook. If your Jupyter

notebook is not installed, you can use

something which is called as Google

Collab, right? Go to Google, type

Collab,

right? Let's say you write Collab,

right? And then you will see this

option. Click on Google Collab and it

will allow you to code in Python, right?

So we are now going to discuss about the

numpy package in python. Okay, numpy

package in python.

So numpy is a

fundamental

package

for data science in Python. Right? It is

one of the most fundamental packages for

practicing data science in Python.

Right? Why is that? Why is so why is

numpy so fundamental? What's is so

special? NumPy package

gives us a new data type

for handling

data in Python

called as

N D arrays, right? ND arrays which

stands for

this stands for

N dimensional

arrays right n dimensional arrays right

this stands for n dimensional arrays

so till now

till now we have studied

about

list integer

tpples,

strings,

right? Out of which

out of which

the data types

such as list

pupils have been used to store data,

right? store data, right?

And range

used for generating

new data which is primarily sequential.

Right?

Now there is one now there is one

problem right? There is one problem and

there should be a question that there

should be a question.

Why do we need a new data type

to work with data science,

right? Why do we need this? Yep. So the

answer to this question people the

answer to this question uh lies in a

small explanation right? Yeah lies in a

small explanation which is that

Python

is a

high

level

language right? Python is a highlevel

language, right? And

a highle language

is usually

very

distant

from

hardware.

A highle language is close to hardware

or distant from the hardware. Did you

not attend the Python programming

essentials?

What is the type of programming language

which is closest to hardware? A

low-level language.

If this is my hardware,

yeah, this is my OS.

This is my application layer. Right? So,

hard level langu language is here and

low-level languages here. Right? Which

is closest. So binary languages,

assembly languages,

right? All these are closest to the

hardware because where is the processing

happening? Where is the processing

happening of the data? At the hardware,

isn't it? Processing of data

is happening

at hardware.

No worry. Actually it is happening at

hardware right? What is processing?

Processing is signals of zeros and ones

right? What are zeros and ones? These

are electric signals.

These are the electric signals right?

Which is basically on and off. Right?

And it is communicated to the hardware

through the help of resistors

and microprocessors.

Right? Isn't it right? Why a computer

only knows zeros and ones? Because zero

is off and one is on which is the

electric current

right electric current to activate or

deactivate certain things right true and

false gates right so it is happening at

hardware so now people if you understand

this part then try to logically connect

it to what I'm going to say when you are

studying data science

what kind of data you'll be dealing with

big data,

right? And as the name suggests, it will

have a lot of volume,

right? It will have a lot of volume

other than a lot of other things, right?

It will be very very big. And on this

large volume of data, you'll be doing

processing.

You'll be doing processing. What does

processing means?

What does processing means? Operations.

So where is this operation happening?

This is happening in hardware

and for hardware which is the closest

language to hardware a low-level

language.

But now people but now we have a

situation in front of us. What is the

situation that what are we trying to do

data science with? What are we trying to

do data science with?

Python, right?

And Python is what?

A high level language,

isn't it? Yeah. So, there's a

discrepancy. Yeah. A big one

because hard level, high level language,

these

are slow

in processing,

right? These are very slow in

processing, right? So for these kind of

languages to handle this kind of data

and these kind of operations yeah is

very difficult right so let's let's keep

let's keep this part aside if you

understand this now let's go to the

second point right

when you learned Python

on a scale of 1 to 10 how easy was it

the ease of use of Python

It's relatively a higher number, right?

Relatively a higher number. So now guys,

when I talk about data science, right?

When I talk about

data science, okay?

Right? When I talk about data science,

my thing is that this will be used by

masses,

right? Will be used by masses. A lot of

people managers, programmers, business

analysts, data analysts, possible right

who are from nontechnical background who

don't know coding they also can do data

science because data science is a

general thing isn't it? Understanding

the data it should not be limited by

your capability to understand the uh

technicalities of a very complex

language. So for these people which

language is suitable which is Python

right? It is easiest to understand. It's

a high level language almost like

English. Yeah. So, Python is a simple

language. So, in this part people,

Python fits the bill, right? Which is a

bigger thing. In the second part, when

we talk about the operations,

we talk about the operations. In this

part, there is a problem, right? Python

fails,

right? Python fails, right? because it's

a highle language. We said that okay

there is a language called as C right

which is a middle level language

right and C language people is used to

create OS operating systems. It is used

to create networks

networking applications.

It is used to create games,

right? All the things which are close to

hardware,

C is used, right? So C fits this bill,

right? C fits this bill.

Python

said that okay, if C fits the bill and

Python is written

in C, written on C, right? It's written

on top of C language. Now what happened

was Python said okay if my intrinsic

data types my intrinsic processing is

not suitable for data science but my

highlevel nature is let's do one thing

let's take C language and use its power

right use its power that it is very

close to hardware and let's create a new

data type right let's create a new data

type which is written on top of C and

which can integrate with Python

seamlessly. And people this new data

type was called as array

and this was given to you by something

called as num py package. Right? Nump py

package. It defined a new data type

which was array. And along with defining

the array, it gave various operations.

Right? It gave various operations

one could

perform

on arrays,

right? One could perform on arrays.

It's simple, right? program the the

power of processing lied with C right it

was lying with C. So we developed a new

data type using C on top of Python and

that new data type was called as array

and this array was defined in a new

module which was called as numpy module

which told you how to create the arrays

and then how to manipulate those arrays

for doing data science. We had options

like Java, we had options like C, right?

We had options like forotron to be used

for data science but we chose Python

because of its simplicity and the simple

syntaxes that people from

non-programming background could also

use Python to do data science.

Right now the limitation was that

because it is slow because of being high

level we needed something which could

make it fast and that was using an

external data type which is not internal

to Python and that was array and this

array is defined inside a new uh module

or a library called as num py which is

numerical python right numerical python.

Back to the programming right where we

have understood right about this

question right.

So

numpy

essentially is

built on top of

C

language which is

compatible

with Python.

It leverages the power of closeness of C

with hardware,

right? C with hardware

which eventually

makes the processing

faster in

Python. Right?

So this is what the first part is right

this is what a first part is right now

the second thing is people so this is

about the performance bit right these

are the performance bit the above

is about the performance

of

right now coming to the next part which

is the memory efficiency Y right memory

efficiency right so

arrays created

by nump py in

python

are less memory

exhaustive

than lists in Python right and I will

prove these points later on to you

through code, right?

In list, right? In list,

each item is an object, right? Each item

is an object, right? I hope you remember

this, guys. Each item is an object,

right?

And it holds, right? It holds

meta information

like

references

and types,

right? Etc., right? A lot of information

it holds, right? This makes

list consume

more memory, right? But but

arrays in

numpy

are contiguous

which means

that they do not

create objects

but rather

directly store

the data

in

continuous

memory

blocks

one after another. Right? Also the

arrays are homogeneous in nature. Right?

You can only store the homogeneous data

in array unlike lists. In list you could

store different different data types.

Right? But in arrays you cannot right.

you have to store the same kind of data

in the array homogeneous. So it's a

contiguous memory block which is meaning

that you can store data in continuity

right one after another in the memory

block. So the access is faster the

memory location and the memory

efficiency is very higher right as

compared to the native data type like

lists or tpples in Python. Arrays

are basically

vectorzed operations

right I'll talk to about talk about to

you with vectors what are vectors right

they are the vectorzed operations and

they are way more convenient

to deal with as compared to list

in Python right in list you have to go

through a lot of loops, right? We saw

that we have to go through the list

comprehension. But you will see in

Python pandas, sorry, in Python numpy,

the vectorzed operations are very very

simple, right? They are very very

simple, right? Again, for all this, I

will give you examples, but uh it will

take some time because you have to

understand what arrays are first, right?

Okay. Then

the most important

other packages

which are pandas,

mattplot lib,

cb bond,

sklearn,

cypy

all are written

on top

of

numpy

package.

That's the reason for this reason

it's called as

fundamental

package.

All the other packages which make your

life easier as a data scientist where

you don't have to worry about code.

All these packages

make our data science

journey smooth

because we have to worry

less about code and more about

logic.

Yes. So for these for the understanding

of these packages it's very important

that we understand numpy first and then

we move forward right

now

then let's get started. So the first

step right the first step which you have

to uh see right the first step which you

have to see while using numpy package is

basically from where will you import the

package right from where will you import

the package.

So to import the package num py we write

import nump py as np right where np

is an alias right it's an alias

right so please do this import nump py

as np

right if you don't get any result for

this then you can write pip install

nump py right pip install nump py and

just execute this right when you will

execute this it will give you this

message or it will give you it will

download this package for you right pip

stands for

python

index package

right

it is basically like play store,

app store,

right? Or Windows store

for Python,

right?

So, you're just going to these Play

Store,

Windows Store, App Store of your Python

and asking them to download this for

you, right? It is also called as the

package manager

right pip

if it is done you can also check the

version you can say np dot

version

and it will give you the version of

numpy

right

numpy is an opensource

package,

right? Yeah, that's about it. Yeah, it's

an open-source package, people,

right? Open source package. And if you

want to see the code, you can go to

GitHub.

Go to Google, write the uh code for

numpy. It will show you it on GitHub.

Right. done. So let's say I okay so I

say

we have when we so okay before this let

me come to a little bit of theory before

I do this with you. So now guys I said

that num py

has arrays

as

data type

right and this is written on C which

runs directly

on hardware

also this numpy array

is basically basically a vector,

right? It's basically a vector. Now,

what is a vector, people? What is a

vector? A vector is a quantity which has

sign

plus magnitude,

right? It has a sign and magnitude. If I

say this is a cartition space, this is

I, this is J. And I say this this is 3 I

and 4 J right 3 I cap 4 Jcap. So this is

a vector right and this is the direction

guys. If you have studied elementary

maths you would know this.

Yes people this is a vector. If I draw

another like this

then this is another vector. So I will

call this say

uh 2 I and 5 J. Yeah, this is another

vector and this is the theta right. This

is the direction.

This is the direction. And what is the

magnitude?

3 I 3ยฒ + 4ยฒ which is 9 + 16 which is 25

under root which is 5. So the magnitude

of vector is five and direction is equal

to theta. This is a vector quantity.

Now what is this? What is this? This is

a scalar.

This is a scalar. Yeah. Only magnitude

isn't it? This is scalar only magnitude.

And then I have 2a 3. Right? This is

what? This is a vector.

It has two dimensions

or one dimension. Only one dimension,

right?

This is one dimension vector.

Yep. One dimension vector. Now if I say

this 2 3 4 5, what is this called? This

is called a matrix,

right? Right? This is called a matrix

which is what collection of vectors

and this collection of m vectors matrix

is called as two-dimensional. Right? It

is called as two-dimensional.

Now if you have this

so these are stacked behind each other.

This is one. This is two. So this is

three right? So we have three layers

in matrix.

So how many dimensions will be this

people?

One dimension, two dimension and three

dimension. This is a threedimension

matrix,

right? Or threedimension vector

or three-dimension array.

I'll repeat once again. What is a single

value? A single value is called as a

scalar. Right? It only has magnitude.

Now when you have multiple values,

right? This is called as a vector. It

has a direction. It has a magnitude. And

this is single dimension. Now multiple

vectors right like this or maybe you can

say like this, right? Are you

understanding why I'm calling it one

dimension? It can be either this

dimension or it can be this dimension.

In any dimension you stack two vectors,

you will get yourself a matrix. Right?

You will get yourself a matrix which is

now two dimensions. It has rows and it

has columns.

Right?

Now if you stack multiple such matrix

one after another, right? This becomes a

threedimensional matrix. And this can go

up to how many dimensions people? How

many dimensions this can go up to? It

can go up to this

can

go up to n dimensions,

right? N dimensions,

right? So now do we get it? Why do we

call it ND

arrays?

Yeah, n dimension arrays,

right? That why are we calling something

what?

Right. N dimension arrays. We are only

capable of viewing three dimensions

people. It can go up to 100 dimensions,

500 dimensions, 1,000 dimensions, any

dimensions.

Scalas are least important.

vectors are more important. So now we

have

multiple

dimensions in

arrays

namely 0D,

1D,

2D, 3D and so on till the N D, right? NV

arrays. Right now let's create our first

array. Right? Let's create our first

array. And this will be a zero

dimension

array, right? Zero dimension array. How

will you create this? You will say a r0

is equal to np dot array and you will

mention what a scalar value is. What is

a scalar value? It is a simple value. I

say two, right? np dot array equal to

two. And when you will now check or

print the type of ar r0, it will tell

you class num py nd array. Right? It is

a zero dimension array. And if you want

to check

the dimension,

you just have to write a r0 dot end div,

right? And it shows you that there is

zero dimensions present. So what is this

in short? This is a scalar. I have

entered a single value people 2 200 500

whatever you want to enter. And the

syntax is np dot array np dot array.

You're instructing nump py package to

fetch the function array on method array

and convert this into that particular

data type right array zero

and when you check the type type is nd

array but what is the dimension of this

nd array this is zero which is nothing

but a scalar right we have created this

this is what we have created

Right

now guys,

if I created this list, right?

Say Lis is equal to

Yeah. So this was this was a list,

right? This was a list and this was this

is what people what is this that we have

just studied according to that? What

dimension is this list?

So what I'm trying to do is I will

create a

one deal

array with a or let's say from a list

right let's let me show you how do we do

that okay

so I say

hurry read the error name ar r not

defined. Why? Because you have created

array from a a r0. Come on hurry.

Right.

I will create a onedimensional array.

How will I do that people? I will say a

ar r1 is equal to np dot array. And can

I pass 1D list inside this? Or can I say

I can pass lis inside this? When I do

this people now see what will happen. A

R R1 will be equal to this right and if

you say let me say print a ar a ar a ar

a ar a ar a ar a ar a ar a ar a ar a ar

r r r r r r r r r r r r r r r r r r r r1

it will be like this okay this is your

a ar r r1 dot nim

you will see that it gives you one right

this is a onedimension array right on

dimension array

Yes. So for example

when I say scalar right when I say

scalar I say

35s right? when I say vector

1D I say

uh

so these are suppose my marks

right now I say 35

40

50 right so now what are these my marks

in three subjects

yeah marks in three subjects this is my

say Hindi this is English and this is

maths right or let's say science because

not everyone has Hindi science English

and maths right guys now if I have to

create a matrix

of two dimension what will I write so

that means this is one student this is

one student people isn't it guys yes no

maybe so now in matrix we will have what

we will have multiple students

Yes. No. Maybe in a matrix people we

will have multiple students. Suppose

this was S1. Now you will have S_sub_1,

S_UB_2, S3, S4 like this.

And each student will have their own

individual list of marks. So can I say

that I'm making a nested list?

Can I say that people? I'm making a

nested list. So now let's make it okay

from a

nested list. Okay, a nested list.

So I'll say ar r2 is equal to np array

list. I will have to create a list

first. Lis2 is equal to. So this is my

first bracket. What is this bracket

representing? this bigger bracket. Now I

will put another bracket inside this and

I will write 1 1 22 33 3. I'll put a

comma again. Write a comma. Then I will

say

4455 666 comma 778899

right I'll do this right now I will say

list to

two

when you will do this you will see that

an array like this has been created

Right?

like this AR R2

right this has been created right when

we check the dimension it is two

dimension array right

yeah marks of three different students

in three different subjects

and this is the same technique you can

create a three-dimension array how will

you create a three-dimension array

people

if I go here how will you create a

threedimension array

Now suppose I have data in this and this

is my master list. Okay, this is my

master list. In this I have data

and this can be represented like this.

This is my

first matrix isn't it? And this is the

vector inside this

V_sub_1, V_sub_2, V3.

Then this can be called as M1. And now

to create a three-dimension setup, how

many M1s do you need? You need multiple

M1s, isn't it? You need M1, M2, M3,

multiple matrix like this people.

like this matrix 1, matrix 2, matrix 3.

So what will you do? You will have you

will have what people?

You will have another yellow, right?

And you will have inside this yellow

multiple purples.

Isn't it

right? Again like this.

Yeah. Like this you will have it people.

So can I say people can I say that as I

am increasing the dimensions as I am

increasing

the dimensions

I am putting 1D sorry 0D right and okay

again in this V_sub_1 in this V_sub1 do

you think you will have multiple scalers

people can I say that

can I say multiple scalers create a

vector multiple vectors create a matrix

And multiple matrix create one

three-dimensional matrix.

Can I say that? Let me talk to you about

an image. Right?

Image, right? What is an image made up

of people?

What is an image made up of?

H

pixels.

Yes or no? No,

not frames. Frames is basically videos.

Pixels are creating an image, right? So,

we have how many pixels here? 1 2 1 2 3

4 5 6 7 8 9 10 11 12 13 14 15 16. Right?

Suppose

this is one vector, right? This is one

vector and this is one scalar

right? Scalar 1, scalar 2, scalar 3,

scalar 4. And this will make vector

v_sub1.

This is v_sub_2, v_ub3, v4. And together

together can I call this m_sub_1 and

call this blue?

So for a colored image, how many

channels are there people? How many

channels are there? What do we call it?

We call it the image as

RGB

RGB image

that means red,

green,

blue.

So this is blue part. So similarly you

will have a red part in front of it.

Then you will have a green part and then

finally you will have a blue part. Yeah.

Are you understanding guys? Why do we

require three-dimensional arrays?

Yes. Suppose now you want to make a

change at this this pixel. So you will

go to the third layer which is the blue

layer. Then you will go to the third

column. You will go to the third column

and third row. And this is how you will

reach this pixel. Everyone? Yes. No.

Maybe.

Yes. So this is the reason why we need

to create a 3D array.

Right. To ingest information like this,

right? To ingest information like this.

So we can create a 3D array also. Right.

Right. And how did I tell you? How many

brackets will I have? Squared brackets.

I will have three squared brackets. So

now this is just one student.

Right. Now I'll put a comma here.

Right? Right, I'll put a comma here and

I will start.

Right, I'll do this. I'll say this list

three

a r3

list three ar r r3

a r3

r

and now you will see that there's a

threedimensional array

right guys

right this is a threedimensional

array

Yep. Moving on. There are multiple ways

create arrays, right? The first one we

have done.

So we have done

from lists.

From list we have done.

Then second will be

from uh we can create a zero array

right we can create

on's array

right then we can create custom array

right I will show this all to you

right I'll show this all to you so let's

start with the on's array sorry zero

those array

right what do you have to do you have to

write so let's say zero

ar r r0 okay a zero dimension zero array

so you say np dot zeros

right and you create

a two

right

a two right and if I say

uh 0

dot end

right you will see it is a onedimension

array

right it is a onedimension array

right so two is by default taken as so

let me just show this to you

it will look like this right it is taken

as horizontal what is the dimension of

this guys a vector what is the dimension

of this vector

no no it's 1. It is basically 2 + 1,

right? 2a 1.

Sorry, 1 comma 2. My bad.

1 comma 2. Isn't it? Now, what if you

had to create a 2 + 1? What? What if you

had to create a 2 + 1? Right? So, let me

just show that to you.

So, I say wait

like this. Okay.

Now when I do this, I say a ar r r once

and I say here

2, 1, right? I say 2a 1. Now you will

see people what will happen to this.

Now

because you have created right specified

two dimensions, right? What will this be

converted to now?

H what will this be converted to? This

will be converted to

a

two-dimension vector. By default, it was

this, right? Which was this vector,

right? By default, it was this vector,

right? What is the dimension of this

vector? 1 + how many values you put

here, right? 1 + 2 like this. So we call

this only one dimension. We call this

only one dimension. But now when I will

run this one, you will see yes 2 + 1.

And now you will see this is

two dimension, right? You see this is

two dimension. Now clear people just the

orientation has changed. But now you see

the brackets there are two brackets now

because what have you now instructed?

You have now instructed Python and

rather numpy to create the vector as 2 +

1. So you have said give me this zero

and give me this zero here. So the

moment you do this you are now

specifying the rows and columns.

Now this 21 right let's say this is 21.

Now let me create a another two cross

two dimension matrix for you. Let me

call this 10

comma 10. Right? So how many rows and

columns will it have people?

How many rows and columns will it have?

I will say 21.

It will have 10 rows and 10 columns like

this. You saw this? Yeah. 10 rows and 10

columns. Right?

By default, it is 1 +2 like this. It is

a one-dimensional vector and this is a

two-dimensional vector. What about a

three-dimensional vector? 0 dot 0

uh a ar r3

is equal to np dot zeros, right? np

do.zer,

right?

H should I just write 3a 3a 3? And I

should check for this.

Yeah, you will have a 3 +3 vector 3 + 3

matrix with three matrix stacked behind

each other. Right? If I say four, this

will be four. So how do we read this?

number of layers,

number of rows and number of columns.

So, can you help me with a syntax? Can

you help me with a syntax which can

create me a 3D matrix of five layers,

three rows and three columns? What will

I write?

Five layers, three rows and three

columns. What will I write?

5 33. Yeah, you'll get five layers. 1 2

3 4 5 right like this. Suppose suppose

you have to represent

an image

with

with say

RGB channel

and

uh say 256 and 256 pixels. How will you

create this? So I'll say img right image

is equal to np dot

zeros and I will say inside this

three channel 256 cross 256

right and when you will just run this

image it will be like this right it will

be like this this is one channel this is

two channel and this is three channel

And uh we understood that numpy is a

fundamental package for data science in

Python. Right? This package gives us a

new data type to work with which is

called as n- dimensional arrays. Right?

Now what are arrays? What are arrays?

Arrays are nothing but vectors right

which are stored in a contiguous block

of memory which means they are stored

continuously one after another and they

do not get converted into the object

unlike the list and they are way faster

they are more memory efficient than list

I will prove this fact to you today with

the help of example through the help of

code right so why numpy arrays because

numpy is a package which is built on top

of C language which is compatible with

python and C being a middle level

language interacts directly with the

hardware. So whatever operation you run

in a fact that is getting directly

executed on the hardware itself. Right?

That is the reason why the uh

performance is way better when we try to

use the n dimensional arrays. Along with

this these syntaxes the type of syntaxes

we used to type uh in list I will show

that to you also today with the help of

example are way simpler when you try to

do them with nd arrays right so arrays

are basically vectorized operations

right and they're also very convenient

to deal with as compared to lists and

other native data types in python right

and the most important part is that in

our data science journey whatever other

packages packages we will use. Right?

Again, what are packages? They have

predefined things stored for you so that

you can leverage them and focus less on

code and more on logic. Right? You need

to be aware about the logic more than

the knowledge of the code. Right? So, we

have packages for that which contains

methods inside them which you can use

and uh without any further calculations

you can work with them directly. Right?

So that is how we started with numpy

right and the syntax to import numpy was

import numpy as np where np was an alias

right. Uh you could do pip install numpy

if someone did not have access to numpy

if numpy was not coming by default. You

can use pip install numpy which is

python index package right and uh this

is like play store app store window for

python. All the packages are stored in

pip and you can call pip you can ask pip

to download that package for you so that

you can use that right it's basically

the package manager so numpy is an open

source and if you want to see the code

you can go to github and check the code

out for yourself right in numpy we have

n dimensional arrays now the question

arises people that why do we need numpy

right so my answer to this particular

question is that when you will deal in

data science right when you will become

a data scientist you will be dealing

with data Right now my question is how

will you ingest how will you make the

machine ingest the data right there has

to be a way right for you to input the

data to the machine right to make

manipulations to the data to read the

data so all this is started as the base

package of numpy arrays right other than

that it becomes very difficult and

cumbersome for us to deal with that and

this provides us a lot of ease and

flexibility to deal with such massive

amounts of data which you will along

with me as we will move forward in this

particular course. Yeah, perfect. Now

people, let me just uh pull up the PBTs.

This is what we're discussing people. We

start with something which is called as

scalar, right? We start with something

which is called as scalar which is a

quantity which only has magnitude.

Right? In numpy language, this is also

called as 0D, right? It is called as

zero dimension. Then people we have

vector right and vector has a constant

dimension. It has only one dimension

right you can interpret it as a row or

you can interpret it as a column it

doesn't really matter because this is

only one single dimension right so

usually it will be written as five comma

blank right there will be nothing

written in front of it so this in numpy

terminology and nomenclature is called

as one dimension right when you move on

then you get combine couple of vectors

you get a shape and now that is called

as a matrix and In numpy terminology it

is called as two-dimension right and

when you try to stack multiple

two-dimension matrices one before with

one after each other or one before each

other then they become something called

as three-dimensional and now in my

capability I don't know what four

dimension looks like but there is a high

possibility that you have n dimensional

data right it has it has n we are

dealing with n dimension data right

suppose with this I also add time right

at t equal to 1 at t=2 that will serve

as the fourth dimension for this data

but how do how does it look like I don't

really know that right because humans

are only capable of visualizing 3D three

dimensions at max right so you can go to

n dimensions and hence the name n

dimensional arrays right nd arrays post

this right post this we moved on to

create certain things and I tried to

explain you the data right so the data

will look to you like this right You

might have a scalar quantity which is

marks right one marks right now if I go

on to vectors in one day it can be marks

of one student

right marks of one student in science

English

and maths right 35

40 and 50 out of say 50 right three

subjects so this will be characterized

as a vector right what will be the

dimension written for For this it will

be 3 comma nothing. This will be zero

right shape will be zero. For this

matrix suppose we have 1 2 3 four

students and each student will have

three marks.

Right? Each student will have three

marks. So what will be the shape of

this? We have four rows and three

columns. Right? So this will be the

shape right of this 2D matrix right and

now if you stack images right one behind

each other then it will be like this

right image I gave you an example so

this has 4 + 4 pixels so the shape will

be 3 + 4 + 4 right this will be the

shape for this particular 3D matrix

right guys so this is how you input the

data just to tell you a little bit more

since generative AI is very popular

these days. Right? So what if I tell you

the fact that the Chad GPT

which you use or you might have used

right has

never seen

a single

word

in its lifetime.

Right? All it sees

is numbers, right? Only numbers. How do

we see numbers? Suppose I say my

name

is

Raghav.

Raghav

is a

nice

name. Right? So these are two data,

right? These are two data points. Now we

all know that computers do not

understand these right computers do not

understand these right there is nothing

no understanding for computers to know

what text is right it only knows 0 and

one yes dhika right it only knows zeros

and ones so see how we will convert this

so there is something called as

vocabulary

right so vocabulary are nothing but the

unique words

right how many unique words do I have in

this my name is Raga four. This is not

unique. This is repeating. This is

repeating. Five, six. And this is

repeating. So I have six words. So now

guys, I will do something called as word

to

right where I will convert these words

into vectors. How will I convert them?

Look at this. So I will have suppose

this is S_sub_1,

this is S_sub_1 and this is S_sub_2,

right? So I will represent

S1 as

right. I will have a vector

of size six. How? I will say 1 0 0 0.

Right? 1 0 0. How many elements does it

have? Six elements. Right? Name will be

0 1 0 0 0.

is will be 0 0 0 1 0 0 0

and ra will be 0 0 0 1 0 0 right this is

s1 my name is raghub now when it comes

to s_ub_2 right when it comes to s_ub_2

how will I enter this s2

0 1 0 0 ragh what is is here

00 0 1 0 0 0

what is

0 0 0 1 0 or nice is 0 0 0 1 and name

will be 0 1 0 0 0 0 right now this will

be the vector representation of these

two sentences just to tell you a fact

GBD3 right GPD3 model right GPD3 or

GPD3.5

they have vocabul vabulary

of 30,000 words, right? 30,000 words.

And each word, right? Each word

has a

dimension

of

12,500

numbers. Right? What do I mean? Suppose

I say Raghub.

So it will be one word and it will be

represented by five 11,500

different numbers

right and like raghub there will be

30,000 words in this GPD model right

30,000 words in this GPD model right and

this total number of parameters which

get trained in the neural network which

we will learn later on neural networks

they are almost close to 1

75

billion

parameters right 1.75 billion parameters

so why I'm telling you all this because

to showcase to you that what is the

importance of vectors right in the

entire machine learning and data science

yeah people so this is the reason people

now my question is how will you create

these vectors how will you read these

vectors Right? The answer is through

numpy package because it is the base

package. Clear people? Yeah. I hope

today's class will be uh interesting for

you because you will know the context.

Why are we doing it? Yeah. So I'll try

to show that to you how we convert

things to vectors.

Right? Okay. Let me go here now. Right.

Let me go here.

So people uh we started using numpy. So

I started with the zero dimension

arrays. Right? Zero dimension arrays. So

zero dimension is nothing but a scalar.

So I created a ar r0 which was np dot

array and I entered a single word single

uh element inside this which is nothing

but a scalar and then I checked the type

of uh ar0 also right and then I check

the dimension also. So the answer was

two class was numpy nd array and the

dimension was zero right exactly what I

had mentioned in my pb

right it will be having a zero dimension

like this right same thing has been

proven

right because it's a scalar now coming

on to one dimension right coming on to

one dimension I create a list which is

nothing but a one-dimension data type

right now I create an array which is a

ar r1 from array from this particular

list lis and then I check the type of

print ar1 check the type of ar1 and the

dimension right so when I execute

right then you will see that it was this

numpy array and the dimension was one

right exactly like this so if I show you

something else say print a ar r r1 one

dot shape

you will see it's 4 comma empty right

and I show it to you here

it will be empty right empty and this

will be comma 1 right so this means that

it has only four elements right if I

increase these elements to say 55 5 66

77

then it will become 7, blank, right?

Which means it has seven elements as a

vector. Now we create something with a

nested list right which is like this. So

with a nest I want one bracket which is

running outside right then inside this I

have one two and three lists inside one

list. Right? So this is a nested list.

This is marks of first student, second

student and the third student. Right? So

I do this and you see it is this right?

And I will show you the shape also

right. It will be 3 + 3 rows and three

columns. Three rows and three columns.

Right? Now similarly we can also create

a

the 3D matrix

right with

two levels right level one and level two

and 3 + 3 so the shape will be what

people can someone guess the shape

what will be the shape of this I've

shown you

Right? If this is 3 4 then what will be

this?

Three rows and three columns. Right? So

when you will execute you will get 2 3 3

right 2 3 3

right

right now people there are multiple ways

right there are multiple ways to create

arrays and we should know them because

all of these comes very very handy. Not

right now. I don't have enough context

to give you right now. But later on you

will see with me or with some other

trainer that how these will be used in

deep learning specifically, right? They

are the key of deep learning algorithms,

right? Where we initialize some weights,

we initialize some biases and those

initializations are nothing but

multi-dimensional numpy arrays, right?

Numpy arrays.

Okay.

Like for example, suppose I have this

I have to multiply this with some random

numbers, right? So how will you generate

these random numbers? You will generate

them through numpy. And you can generate

them in a specific kind of uh shape,

right? Which is 2 + 3. And then you can

multiply them. You can multiply the

matrices and you can get your output for

yourself. Right? So this is the way they

are used.

So we saw the first thing from list we

have already covered. Then now we are

moving on to creating zero arrays right.

So I create a zero array of one

dimension right of one dimension which

is 0 0. They are represented in floats

right. They are represented in floats

0.0 zero. Right? Now you can create a

two-dimensional zero array. Right? You

can create a two-dimensional zero array

which is you have to mention just the

shape inside 2 + 1. So it will have two

rows and one columns, right? Two rows

and one columns. The difference here is

the difference here is that these are

one dimension and these are two

dimensions. Right? You have explicitly

mentioned the rows and columns. So you

can expand this to 10 + 10 also.

Right? You can expand in 10 + 10 or 10 +

6 whatever you feel like yourself.

Right? It will have 10 rows and six

columns. Right?

Now you can also create

the 3D arrays 3D zero arrays.

Right?

Which is 5a 3a 3. What does five means?

What does five means? First element

represents the number of layers. So you

have five layers, right? It's five layer

deep. Then you have three rows and three

columns, right? So it will look

something like this.

1

2

1 2 3 4 5 right like this something like

this right it will look like this tab 1

2 3 1 2 3 1 2 3 right 5 33 okay yes the

number of matrices hurry what I

represented

right layers

rows

columns right layer rows and columns.

Clear?

So now when you execute this, you will

get an arrangement like this. Okay. I

try to show you this thing with another

example, right? Which is I created an

image of an RGB image of 256 cross 256

pixels, right? Which have all zeros

inside them. And this is how it was

created, right? Three layers RGB 256

256. So this is how the image will look

like

right.

This is how it will look like.

Now you can also create

arrays with ones. Right? Exactly the

same way you created it with zeros. I'll

give it to you. I'll give you 5 minutes

time to create them. I'll show you one.

So I say

a1 is equal to np dot

once

and inside I pass

two

I check a1. So this is a array like this

right? It is an array like this. Now you

create create two dimension

and

three dimension

arrays of one. It is basically

as a float. Hurry. It's represented as a

float. Right. It's represented as a

float. Okay.

Right.

Perfect. Right. Also guys uh with this

right also with this you can create the

custom arrays. Right. You can create the

custom arrays. Right. How do we create

custom arrays people? How do we create

the custom arrays?

You have created now zeros. You have

created now ones. Now what is left that

you create the custom arrays. Uh forget

about this. I will come to this later

on.

Right? Let's create

custom arrays. Right? So the syntax

remains the same. Right? I'll say cus

arr is equal to np.

Right? This is the syntax np.

Right? And you will say 6 + 6 and

suppose you want an array of all fours.

Right? This is the dimension 6 + 6. And

this value after comma is basically the

value which you want. You execute this

and you copy this paste this and you

will get the arrays of fours for

yourself. Right? If you want of 10, you

will get of 10. If you want 10.3,

you will get 10.3. Right? anything which

you want. If you want case, you will get

case, right? All the examples. So, let

me just show that to you.

Yep. Like this. Now guys, how did we

create

or how did we use

range in Python?

Can you use

range to generate

numbers between 20 to 50,

right? 20 to 50. Can you give me the

syntax quickly? How did you do that in

range?

How do we do that? We said R is equal to

range

20 to 51. Right? And then I said

I in R

print I,

right? And this is how I got the

numbers, right? So similar

to range in Python,

we have

a range in num py. Right? How do we use

a range? I say

uh a range

ar r is equal to np dot arange. Right?

And then same syntax I will say 20 to

51. Right? 20 to 51. And that's it. And

when I will check my AR range error, you

will see I have generated myself numbers

between 20 to 50 and a range in numpy.

Right? A range in numpy. Yes, if you

want a interval so you can use this say

a range one and after comma you pass the

third argument. Suppose it's three. So

now it will jump three times, right? 20

23 26 29 32 35 like this up till 50.

Now guys there is something which is

called as lind space

right. What is lindspace?

It stands for

linear spacing

which means

between two given numbers.

This function will fit the required

number of

numbers. Right? For example, suppose for

example,

we need to create an interval

from 0 to 1. People, there are infinite

numbers I can have between 0 to 1. Isn't

it?

Infinite numbers I can have between 0 to

1. 0.0000000000001

0 0 1 0 1 01 right I can go in the

infinite manner right now for example

you need to create numbers between 0 to

10 right and you want to create and want

to have

10 numbers in it right so how will you

do this it's not float it's about the

number theory right between 0 and one

you have infinite finite numbers, right?

So you say lindspace is equal to np dot

lindspace, right? np.tlind space. You

mention from 0 to 10, you want to have

10 numbers, right? And when you will

create lindspace,

you will see that these are the numbers

are there which have been created,

right? These are the numbers which have

been created,

right? Nine numbers. Now I say 100

numbers. I want evenly spaced 100

numbers, right? Evenly spaced 100

numbers. How are they even? You can

simply subtract one number from another

and the difference for all the numbers

will be exactly the same. 0.01 0 1 01.

Subtract any two numbers. It will be

0.01 0 1 01. Right? Where do we need

this? We need this to plot the axises.

Right? When you will plot graphs, you

will need access between this interval.

You need five values that works like

this. Okay? Suppose you want from 0 to

10 five different values, right? You

will have five different values like

this, right? Between the gap of 0.5,

right? If I say 1 to 10, you will have

values like this,

right? Like this. Clear? 100 values like

this. Yeah. Clear guys. How do we use

lin space? Suppose you want from 0 to

10. Interval from 0 to 10 and 10 will be

included. Zero will not be included.

Right? We'll start from one. So I go to

one it will be from

one. Why is 0 not included then? Yeah. 0

is included. Right? 0 is also included

and 10 is also included. 100 numbers

between them. Right?

Clear? This is what lin space is. Now

guys, now

suppose you

lohan l space is basically used if you

want to create n numbers between the

range of numbers right between 0 to 1

right suppose between 0 to 1 you are

trying to plot a graph okay and your

values are 0.2 0.3 0.6 six right and you

want to draw a graph so you will have to

mark the axis right the x axis and the

y- axis so you can use lindspace there

and what will it do it will take the

range it will take the interval in

between you want to add the equal space

numbers and then the third argument here

will be that how many numbers do you

want between them so this syntax tells

you that

from

0 to 1

give me 100 numbers. How are these

numbers? Equally

spaced

numbers. Equally spaced numbers, right?

So when you will execute this, you will

see that all there are 100 numbers which

have been generated, right? 100 numbers.

And all the numbers are equidistant from

each other because difference of every

single number from the next number is

0.01

01.

Yep, that's what it does. Right

now guys, now suppose we want to

generate

random numbers, right? We want to

generate random numbers, right? Now

we're interested in generating random

numbers. So we have something called as

random

dot random right. What will it do?

Random.random

will generate

random

float numbers

between 0 to 1. Random float numbers

between 0 to 1. How will this happen?

You will say

rand rand is equal to np do. random dot

random and inside you will mention what

is the dimension that you seek. Suppose

I want 6 + 6. So when you will check

this you will have all numbers for 6 + 6

dimension right 6 + 6 matrix right now

every time you rerun this the numbers

will change because all of these are

random numbers

right all of these are random numbers

now I say 100 multiplied by rand rand

you will see all of them all these

numbers will be multiplied by 00 right

all of them

in one shopping.

Now just like this we can also create

random integers right how will we create

random integers guys

I say rand intore

rand is equal to np dot random dot rand

right and here you will specify that

what is the range of numbers you want

from so I say between 20 to 25 I need

random numbers and then I want it from

in a 3 + 3 format. Right? And now when

you will check your random you will get

random numbers generated like this.

Okay? Random numbers generated like

this.

Right?

If you say 3 + 3 + 3 you will get a 3 +

3 + 3 matrix. Even if you will only say

3, you will get a 1D.

Right guys? You can change the

dimension. So this is the range from

which you want to choose the random

numbers and this is the dimension you

want this matrix or vector to be in.

Now guys, we'll move on to the next part

which is basically properties and again

there are a lot of operations you'll

have to see it yourself right

properties and

attributes

of numpy

arrays right property and attributes of

numpy arrays.

Okay. Now guys, the first one in this

scheme of things is shape of array,

right? Shape of array. What is shape of

array? It tells you

the

dimensions of the array

stored in a

tle. Right?

For example, I say a ar r r r r r r r r

r r r r r r r r r r r r0

right a r r r r r r r r r r r r r r r r

r r r r r 1 a ar a ar a ar a ar a ar a

ar a ar a ar a ar a ar a r r r r r r r r

r r r r r r r r r r r r r 2 a ar r r3

right and then I say

print this

dot shape

right like this and you will see that it

will give you the shape of each array

Right? 0D, 1D, 2D and 3D. Right people?

Shape of the array.

Please try it out. We have used it one

or two times. But this is what shape of

array actually means.

Second is people

end

right is end

right end of array right it tells you

the rank of the array whether it's one

dimensional two dimensional zero

dimensional threedimensional four

dimensional

so again I will do the same and

right I'll say end and you will see it

will give you 0 1 2 3 zero dimension

zero rank one rank two rank and three

rank and it can go all the way up to end

rank

before this I should have also

printed these arrays

Right. These are the arrays.

Yep.

These are the arrays which we have and

these are the subsequent things, right?

Rank copy array. Then guys, the third

thing is the size of

array, right? It tells you

the number of elements inside. How many

elements do we have inside this array in

a ar r1? How many elements do we have? 1

2 3 4 5 6 7. How many elements do we

have in this 2 + 2 ar2? 1 2 3 4 5 6 7 8

9 which is rows multiplied by columns 3

* 3 right and how many elements do we

have in this 3D which is 2 * 3 * 3 which

is 18 right so now when you copy this

right you can use this

and say

size right it will say 17 918

Yep.

Fourth is people.

The D type

of array tells you the data type of the

array. Right? And I've told you we place

only

homogeneous

data in array right what will happen if

we don't do this I will show that to you

also right so we do this right and we

say

retype

and you will see in 64 all of them are

integers right all of them are integers

that's the reason we are getting in 64

right suppose I create a new array a ar

r new right let me say head

right hetro

heterogenous

and I say it is like n dot array

let's say like this

right like this now people when you will

Check

ar r

dot d type you will see it will give you

float just because of one floating point

number inside this entire array it gets

converted to float right between all the

integers if you put one float then it

will be taking float directly right now

let me just copy this and let's say

heterogenous one and let me add another

value which is string and say rather

right and when you will execute this it

will give you U32 U32 here is

representing strings right it is all

called as objects right these are all

string values right so precedences

strings greatest then float and then

your uh integers right if you place the

heterogenous data inside the numpy array

right you only need to put homogeneous

data in the array

now fifth is the item size

of array. Right? What is item size?

It gives you

the bite

occupied

by each element of an array. Right?

Because we assume that elements will be

homogeneous. It will give you the bite

occupied by each element of the array.

Only one element. Okay? So how will it

happen? So let's say a ar r r0

or let's say ar r r1 dot item size

right and you will get eight right. So

why eight? Because

because each data point right each data

point is occupying

the result is 8 bytes. Let me put this

here.

Right. Eight bytes

because each element in ARR1 is

occupying

64 bits which are

equivalent to which are equivalent to 8

bytes. Right? one bite is equal to 8

bits. So 64 bits will be equal to 8

bytes. Right? That is how it is giving

you the result. Now if you are

interested in knowing the entire bytes

right entire bytes then you say n bytes

will give you the

total bytes

occupied

by the elements of the array. Right? You

say print

a ar r1 dot

n bytes right and write

bytes it will be 56 bytes

right why because how many elements do

we have in our ar r1 1 2 3 4 5 6 7 right

7 8 are 56 right 56 total bytes are

being occupied with by ar r1 one. Now

guys, the seventh one

is

as type right

in array.

This will help you change the data type

of the array. Right? Change the data

type of the array. Suppose I have a arr

type which is uh so I'll say print

d type right this is end 64 right and

now what I do is I say print

uh wait let me give you a structured way

print a r1 right let me say

array

R1

right D type of array.

Now

a ar r2 sorry a ar a ar a ar a ar a ar a

ar a ar a ar a ar a ar a r r r r r r r r

r r r r r r r r r r r r r1 is equal to a

ar r1

dot as type right dot as type and let's

say I want to convert this in np dot

right np dot

uh

int 32 right I want to create convert

this in uh ar r r int 32 right when I do

this and now when I will copy these same

things you will see for yourself. Right?

Now the D type was int 64 and now the DT

type is int 32. Right?

If I want I can do this conversion in

float also

float 64. Right? And then I will just

copy this

and I will paste it here.

Right? And now you will see now the

floating point has been activated.

Right? It has been now activated. We can

go till int. We can go till int 8.

Right?

I can go to 16.

Right? And I can do this.

I can go to int

8 also.

Yeah, like this in date also 64 32.

So first one has to be

Yeah.

32

whatever right like this okay you can

convert this

also people this is later on conversion

you can define the data type of the

array while creation time also. How

would you do that? Suppose you are

creating a ar r11 and you say np dot

array right and suppose you take it from

a list and then you just put a comma and

say d type. So what will be the default

data type here people? If I just do this

if I just execute this what will be the

default data type? Int 64 is the default

isn't it? But now suppose I want to

change it right here. I say data type is

equal to float 32 right sorry float 64

np dot

sorry my bad

float 64 right and I say a ar r r11 and

this will be float 64 if you want float

32 it will also become float 32 right

right here while you define

Instead of using as type, you can do it

right here. Right? These things will

come in very handy people because you

will have to save memory because when

your data becomes very very big, you

will be always in a crunch for memory

like this. You want integers, then

integers will be like this

random.randent

like this. Suppose you want to generate

0 to six, right? And suppose you want to

generate

say 100 numbers like this 0 to 6 the

scores

1 to six like this randomly

right suppose you want to generate 100

scores for five different batsmen

randomly it will be like this batsman

number one batsman number to bat number

three, fourth and fifth. Right.

Yep.

Understand the data guys. Now it's the

time to understand the data.

Yep. Now guys, we have methods in numpy

arrays, right? Methods in numpy arrays.

So what are these methods? Now the first

method we have to learn is called as

reshape right. Reshape.

Yeah. So reshape is you can use this to

create

you can use this to create

a new shape of the array. Very very

powerful guys. Very powerful. One of the

most powerful methods in numpy is uh the

reshape right and how do we use reshape

suppose

we have a 1D array of 20 elements right

now to reshape this

reshape it we need to find the factors

Right. Factors of 20. They are what? 1

20 4 5

2 10.

Right. The other factors.

The other factors. Now see what will I

do. Right. Now see what will I do. Let

me create a say random array. Right?

Random array. I say random

arr is equal to entprandom

dot rand right and let me say I want to

create it from 1 to 50 right and I want

it to be having 20 elements right so I

say random arrand

values inside this right random 20

values now see Now reshape

first

I will reshape in 1 + 20 right 1A 20 how

will I do that you just have to write

uh

print

a ar r r sorry sorry random dot ar r

random ar

dot reshape dot reshape and you just

pass in the dimension I say 1 20

right and when you will do this you will

see it is coming now in 1 20 format

right so let me just also write print

right 2D

1A 20

right and now what I'll do is I'll copy

this and I will paste this and say 20

comma 1 right you will see it will be

like this 20 comma 1 immediately with

reshape right let me copy this

let me say 2D I'm still at 2D let me say

2 10 right and you to see this is 2A 10.

Now I can reshape it in

10 2

right 10 2 right then I can reshape the

same thing

in

4A 5 and I can reshape this in 5A 4

right like this guys are you able to see

the power one dimension I'm able to

create two dimensions

And now I will take it a step further

and I will write it in three dimensions.

Right? How will I write it in three

dimension? Let me say this 1 comma

2a 10. Right? This is will also be three

dimensions. Let's sorry 2a 2a 5.

Let me say this. And now you will see I

can have this in three dimensions.

Right?

Yep. I can also say in three dimensions

like this. I want to have five layers

with two rows and two columns. Right? So

you will have it like this also. Right?

So this is how people we can reshape the

array. Very powerful. Very very

powerful.

Right? Very very powerful.

Right? And you can take this

and save print

random dot this and you can say print

dot shape.

Right? So this was the first one.

Yep. like this.

Also people if you want to visualize we

can also go this route.

We can have 10 comma 2 comma 1, right?

It will look like this, right? 10

layers. 10 layers you can have,

right? 10 layers you can have.

Great. Now, second method which we have

to learn is called as

transpose,

right? Transpose method. Right? What

does that do? It interchanges the

dimensions

like rows and columns, right? Yes.

Absolutely. Right. Suppose you have a

matrix, right? Which is

22, 33, 44, 55, 66, 77, right? This is

a. So now when you will a transpose it,

right? The dimension right now the shape

right now is 3A 2. Now this will become

2a 3. And how this will happen? Rows

will now become columns and columns will

now become rows. Right? So let's make

column the rows. Right? Sorry columns

the rows. It will be 22 44 66

33 55 77. Right? So people in transpose

no information is lost. It is just a

change in the view right which is

happening right?

22 44 66 33 55 77.

Why do we need transposition? Suppose we

have two matrix.

This is 11th class mathematics. Right?

one has

a dimension of n cross m and the second

has dimension of a cross b. If you

want to multiply

these two matrix say

M_sub_1 and M_sub_2,

right? There needs to be a satisfaction

of condition. M should be equal to A.

Right? M should be equal to A. Right? M

should be equal to A. This should be

equal to this and the resultant vector

the resultant matrix which you will get

will be of n crossb dimension right. So

often times suppose this is n cross m

this is n cross m right m cross n and

you know that m is equal to a. So what

will you do? You will transpose this

matrix right? You will transpose this

matrix then it will become n cross m and

then m can be equivalent to a. Right?

For example, what I'm saying, we have

one matrix which is 2 + 3 and this

matrix is 2 + 5, right? So, can you

multiply these matrix people? Is 3 equal

to 2? The answer is no. Right? The

answer is no. So, what will you do? You

will just transpose this and this will

become 3 + 2 and this is 2 + 5. And now

you can multiply this and the resultant

will become 3 + 5 matrix. Right? So for

operations like these we need

transposition. Right? So how do we

transpose it?

How do we transpose it? So let's say

again

uh a ar r2 right this is a ar r2 and I

want to transpose it. So I say a r2 t is

equal to uh np.transpose transpose

sorry a r2 dot

transpose

right

and now when you will see ar r2 ts you

will see rows and columns have

interchanged right rows and columns have

interchanged with each other

or let me give you one more example

Uh if this is not clear, let me pick up

this again

right now. I say

dot reshape

into say

2 + 10. Right? So this is 2 + 10. And

now when you want to transpose this so

I'll say this t is equal to this dot

transpose

t or a this and we can check this now

and it will be this. Sorry guys. So this

is going to be it will be like this

right transposed.

And now if you want to see this,

this was the original shape, right? Rows

and columns have now been interchanged,

transposed with each other.

Now guys, the third method,

the third method which is there with us

is called as flatten, right? It is

called as flatten.

Right? What does flatten do? It reduces

the dimension to one dimension. Right?

Any dimension you have, it reduces it to

one dimension. For example, I have this,

right? And now when I say

this

dot_f,

this will become this dot platin.

And when you will check this up, you

will see that it has now become one

dimension. No matter how many dimensions

you have, it will become one dimension.

Right? Let me take this again to show

you one more example.

Right?

And here I say dot reshape into

uh

uh 5 + 2 + 2 right I do this my random

ar r is this right it has five layers

two rows and two columns right so now I

say this

flatten is equal to this dot flatten

and If you will check it now again, you

will see it has now flattened it out.

Yep.

Now why do we need this? We need this

for a lot of statistical operations. We

need this to feed the data into the

algorithms. Right? As we will move

forward, you will understand the use of

flattening.

Right? Now guys, moving on and uh as

discussed, let me now show you

the power of

numpy,

right? Numpy

over

lists and other data types, right? I

will not take a lot of examples. Just a

second, guys.

Yeah. Okay.

Now guys, I told you that

less

take up

much more memory

as compared

to numpy arrays. Right? And I'm going to

prove this to you. Right? Now let me use

let me create a random sequence of

random numbers using range in Python

right and say I create range of 10,000

numbers right range of 10,000 numbers so

what will this give me this will give me

numbers from 0 to 99999 right continuous

numbers right so this is range I will

use

a range in

numpy Y to create

a similar

series of numbers

right so let's say array is equal to np

dot arange

right same thing same done by both right

I've shown you above also now let me

import sis

library right sis package and I will use

something called as get size of right

get size of. What does this do? Get size

of it's a method

which calculates

the bytes

occupied

by a single

element in

vanilla Python. What is vanilla Python?

It is the traditional Python,

right? vanilla Python.

So let me just show that to you. I'll

use this and I will say print. Now guys,

if I get

size of any random number from this

range, right? Any random number. Say I

get size of five, right? And I then

multiply that byes with the length of

random, right? With the length of random

this rand, right?

Right? With the length of random, do you

think I will get the bytes for the

entire

data structure? What am I saying is

suppose

uh I used range

five. So what will this give me? 0 1 2 3

4. Right? This will be the output. So

now I say get

size of say I say two. Right? So suppose

2 is x and then I multiply this with the

length of this series which is five. So

do you think I will get 5x which will

represent the number of bytes occupied

by the entire data type. Anything

randomly any random number this can be

three right? Why not hurry?

Why not?

All of these are integers. So integers

all of 64 bits assuming. So if you

calculate the side of size of this and

if you multiply with the total number of

numbers you will get the total size

isn't it?

Huh? Index is in

no no it's not about that it's about the

element right homogeneous elements

inside this.

I am saying when you use range five what

is going to be the output? 0 1 2 3 4

right now all these are elements

elements of range.

Right? All of them are elements of

range. Right? Now I'm saying if I fetch

the size of one element and multiply it

with the length of the entire range,

will I get the bytes occupied by the

entire range? For example, if I do this,

right? If I do this,

this is 28, right? 28 bytes people. 28

bytes

bytes are occupied

by one element of range right

one element of

range right now if I just I'm saying I'm

just saying if I multiply to find how

many numbers range has

how many numbers

range has

equal to 10,000

right so total

memory occupied

will be will be how much it will be 28

ult*lied by 10,000 which will be equal

to 28,000

yeah and how will you find this you will

say Print

this multiplied by length of RAM,

right? 28,000 bytes. Clear? Now, yes.

Now, this is for the range. Now, let me

use another thing. So, how will you

calculate the length of this array? What

what property and attribute will you use

people?

N bytes, right? N bytes will give you

total bytes occupied by the elements of

the array. Right? We will use n bytes

here. So I come back down and I say

using n bytes for arrays. Right? And you

will see what the result comes. Print

uh array dot n bytes. Right? And I say

bytes. Are you ready to see the result?

Do you see what has happened?

How many bytes this was taking? It was

taking 28,000 bytes. How many bytes this

is taking? This is taking 80,000 bytes.

This was taking 2 lakh 80,000. This is

taking 80,000. Two lakh extra bytes of

memory is taken by range.

Ran is range.

Yep. And if I just go to million

numbers,

right? Million numbers in both.

See the difference it becomes,

right? This is now 3 three 28 million

bytes it is taking and it is taking 8

million bytes. 20 million extra bytes

are occupied right now people do you

believe me? Yeah, that numpy wy are way

more efficient in memory management as

compared to the traditional data types

of Python. Yes. Okay, that's the first

part. Now second is people performance,

right? Performance. So what I'm going to

do is what I'm going to do is I am going

to

import

time, right? It's a it's a module in

Python, right? Suppose I say x is equal

to range

this much right. Okay. And then I have y

is equal to range say

this

to

this. Right? Both of them will have

equal amount of numbers. Same numbers

both of them will have. Right?

This will have say

uh 1 2 3 1 2 3 10 million values. 10

million values. This will also have 10

million values,

right? Both of them will have 10 million

values. Now what I'm trying to do is I

want to add them up right by bit by bit.

I want to add them up right. I want to

add first element of this to first

element of this. second of this to

second of this, third of this to third

of this like this. Okay, I want to do

this. Now what I'll do is I will run a

counter, right? I will run a counter

which is the start time,

right? And this is given by time dot

time, right? Which will give you the

this will give you the

current time, right? After this I will

run the operation. I will say C is equal

to X + Y

for X Y

in zip

X Y right in zip X Y right add X + Y bit

by bit element by element for X and Y in

zip zip is a function right which allows

you to do this operation sequentially

right sequentially right add the

elements of X and Y

element by element right element by

element

right element by element right and then

I'm going to print right so start time

will start and now I will say time dot

time which is now the end time minus

start time so this will give me the Time

taken for execution isn't it guys

will give me

delta of time which is equal to time

taken for operation

seconds

right these many seconds will be taken

right so let me run this and it takes

around say

4.3 seconds right to do this right 4.3

seconds now Guys, see what happens. You

had to write this complex syntax in the

traditional Python. Now let me show this

on arrays. Right? What will happen on

arrays? I will say a is equal to np dot

a range.

Right? And inside a range I will pass

the same values what I have taken above.

Right? And I will say b is equal to np

dot

a range and I will pass the same values

inside

right exactly the same now what I'll do

is I will say same thing

right just I will change the execution

of C will now simply become people A + B

what is simple this or this

this or this

two right do you see the power if not I

will show this to you again right later

on and let me run this and you see the

difference now let me just increase a

couple of zeros right a couple of zeros

two zeros I'm increasing in both the use

cases

it is going on and on right let's see

See how much time it will take

to add say 2 million 1 billion numbers.

1 billion numbers I have asked my system

to add and I want to see how much time

it takes.

Running running running.

Yep. Colonel has died. Kernel has died.

People,

I'll have to restart.

Right. I will have to import

numpy

as np. So let me just remove one zero

from both.

It is taking 4 seconds. Removing one

zero from here also. Right?

And when I do this it takes 1 second. Do

you see guys what is the difference in

performance also right for both of

these? Yeah.

And if you didn't understand this, let

me give you an example.

Range five. This is 5 to 10, right?

And this is basically

adding elements,

right? 5 7 9 11 13 Right. So this will

be what will the output of this? This

will be 0 1 2 3 4 and this will be

output what 5

6 7 8 9 right so 0 + 5 5 6 + 1 7 7 + 2 9

8 + 3 11 9 + 4 13 and the same thing if

I do here

then what will happen I say 5 I say 5

and 10 right and I say C is equal to a +

b and I say c. Same thing you get here.

Right? We can move on. Right? The next

bit guys which we have to understand the

next bit which we have to understand is

called as the indexing in numpy arrays.

Right? Indexing in numpy arrays. Right?

How do we index the elements? Right? How

do we index the elements?

indexing in

nump py

arrays right indexing in numpy arrays

so again you know indexing from basic

python so let's start with 1d for 1

arrays right I will use ar r r1

yeah this is a ar r1 now right this is a

ar r1 Okay.

Now people what I want to do is what I

want to do is I want to fetch right I

want to fetch right you can slice and

dice let's say dice

33 right so again as per our normal

indexing of list what is 33

what is the index of 33 people

two so you will say the same thing print

right A R R1 squared bracket 2 and you

will get 33 for yourself. Right? If you

wish to slice same things, right? 33 to

say 66. What is the index?

33 is 2. 2. Which one? 3 4 5 and 6.

Right? We will write 3 to six. Not five.

Hurry. Five is not included. Remember?

We print

a ar r r1

2 is to 6 and you will get 33 44 55 66.

Right? Simple indexing. Please try it

out. Please try it out. And if you want

you can have this code also. You can

write this code. You will always have

clarity that why do we use it.

Moving on people. Moving on. Let's see

indexing.

in a 2D array. Right? And before I

explain this to you, let me take you

here. Right? So a 2D array will be what?

Right? This is a 2D array. It has three

rows and three columns. Right? Rows

columns. So now for rows indexing will

start from zero. So if you have to pitch

this particular row, right? So what will

be the index? It will be row 0. If you

have to fetch this particular row, the

index will be one. And if you have to

fetch this particular row, this the

index will be two. Similarly, for

column, if you have to fetch this

particular column, right, you will have

column is equal to zero. This particular

column, column equal to 1. This

particular column, column equal to two.

Right? indexing will start from minus

one again right 012

so let me just show that to you right

let's say I call ar r r2 this is my a r2

now

indexing

first

row right how will I do that I will say

print a ar r r2 and I will write how how

will I write this

h I will Write row 0, right? Row 0. So

what will this give me? This will give

me this, right? The syntax is

row space column. Right? So now if you

just pass one, it will give you only

rows, right?

row one, row two, row three. Right?

Similarly,

if you want to create it for columns,

right? What will you say?

Sorry,

uh

columns.

Uh

uh it was

zero

column 1 column 2

column 3. Right?

Right.

So what is my column 1? 76 89 98 76 89

98 right and if you want column 2

uh sorry

if you want column two this is the

column two and this is column 3 right

guys 90 999 99 independently

now if I ask you to fetch me a

particular element

element, right? Say I want you to fetch

me 78, right? How will you fetch 78? You

will say print a ar r2. What is the row

for 78? Which row does it belong to?

012. To which row 78 belongs? One row.

Which column it belongs to? 012.

1. Hurry. Check again whether it belongs

to zero column, first column or second

column.

Right? And you will get uh sorry

01. So this is two. Right? 78. Right?

Second row first column. Right? Second

row first column.

How to define it as a

right? This is your matrix. Right? Now,

what are the index positions for this?

This row is zero. Row, first row, second

row. This is your zeroth column, first

column, second column. Yeah.

Yes. No. Maybe. Are we understanding

this? This much is clear. The indexing

of rows and columns. Now, now if you

have two, so the syntax is

syntax is say this is matrix A. So you

will say A in this A matrix you will

write row,

column. Right? Suppose I want to access

only zeroth row. So there will be no

column. So what will you get? Row number

zero.

Right? Row number zero. And you will

just put it like this. Or you can put a

comma and put colon. Colon means what?

Take everything. So I want all three

columns together. Zero row and all three

columns. So this will be your

show. Let me show that to you.

Right? See this zero row and all the

columns. So what will be your answer? 76

88 90. Right? Then if you want to access

the second row 89 90 99 89 90 99 one and

like this. Clear? Now similarly for

columns what will happen? You will take

all the rows

comma which column do you want? If you

say two what will be the result for

this? What numbers will you get for

this? Colon, 2, you will get 33 66 99.

Now coming to the element, right?

Suppose now you want to fetch 55s,

right? So what will you write? You will

write which row does it belong? Follow

the syntax. It belongs to the first row.

Which column does it belong to?

First column. So what will you get? 55.

Come to come come to this example. Now

you want 78 in this particular array or

matrix. Right? Where is 78? Which row

does it belong to?

Is it first row? Check carefully.

Zero row. First row. Second row.

Yeah. So I put two here. Now which

column does it belong to? First column.

Second column. Sorry. zero column, first

column, second column belongs to the

first column. So I put one here. So when

you put this syntax, you will get 78.

Clear? Now the third thing is slicing

through the array. Right? Now for this I

want I want 90 99 78 99. That means

what? I want this, this, this, and this.

Let me take you to the PPT first. Now,

what I'm asking you to fetch me? I'm

asking you to fetch me these four

numbers. Right? These four numbers. So,

here what will you write? Which rows are

included in this people? Which rows are

included in this?

Row one to all.

Right. One to all.

Yeah. Not two. One to all.

Right? And you leave everything like

this. If you have the last row, you

leave it empty after the colon, comma.

Which columns do you want for this? One

and two. Right? And when these will

intersect, when these will intersect,

you will get this area. You will get

this shaded area, green shaded area. So

I will say I need from column one to all

the columns. Right? So what will this

fetch you? What will this fetch you?

This will fetch you 44

55 66

and 77 88 99. Right? What will this

fetch you? This will fetch you 22 55 88

33 66 99. What are the commonalities

between both of these

access

both of these slices? What are the

commonalities? It is only this much

right

55 66 88 99 right so do you think you

will get your result yeah let's check it

here right I say print a ar r2 right and

this I say

I need row zero sorry row one to empty

and then column one empty and you will

get 90 99 78 90 indexing.

Yeah.

Now guys, let's check it for

three dimensions, right? 3D.

I say AR R3, right? This is my AR r3,

right? So, let me just put an example

here. Suppose my arrays are 1 1 22 2 3 4

4 5 6 7 88 91. Right? This is my first.

Then behind this I have another matrix

which is uh

111 222 333

444 555 666

777 888 9999.

Right. And in my

third one, I have 11 1 1 222 33 33 3 4

44 44 44 44 44 44 44 44 44 44 44 44 44 4

55 555 66 66 66 66 66 6 7 8

9

is that is it qualifying for a 3D array?

Can you access anything on this

particular array if given a chance

separately?

like this one. Now people, if you want

to move between layers, right? If you

want to move between layers, do you

think I have told you one particular

access which is what? Which is the

layers.

So what number will be given to this

layer? Layer number zero,

layer number one

and layer number two. So now if you have

to access this 555

which layer you will go to first you

will go to first layer. Then which row

will you go to? You will go to first row

and first column. So if you pass this

syntax what will you get? You will get 5

five5.

Right? Problem solved. The only

bottleneck was the layers part and you

have additional parameter or argument

for this layer.

So if I come back to my example and if

you have to access this 555 right how

will you do that? I say print a ar r2 a

r3 right and this I say and in this I

say what I have to access this 555 which

which layer number is this? This is

layer number zero. Right? This is layer

number zero. And this is layer number

one. So I say layer number one. Which

row is this

in this particular layer? It is layer

number. Sorry, it is row number one and

column number one. And when you do this,

you will get 5x5.

Right? If you want 999, what you will

do? You will change this to 2, 2, right?

If you want 777 or let's say if you want

98 what will you do for 98

layer number zero row number two column

number zero then you will get this 98

clear guys? Yep. Layer, row, column. In

the same way, you will slice it. Right.

You will slice it.

Yep.

Right. I'll write the syntax

so that you don't get confused. It is

array square bracket. Layer row column.

Right? Layer

row

column.

Right. This is the syntax.

Perfect.

Right. Great. We can also perform some

operations, right? Plus, minus,

multiplication, division between two

arrays very very easily, right? It

should not pose any problem to us,

right? We can do all the operations

which we want to, right? Between two

arrays, right? For example, right? You

had you have

right operations on arrays.

So let's say uh

let's see as a list right list. So we

have list is equal to 1 2 3 4 5 right

now suppose you want to square

each element of the list. Right? What

will you do? You will say s sq l is

equal to x to the power of 2 for x in

l right and then you will say print xq l

and you will get the squared of list

Right?

Original list squared list. Now

in array what will happen? Suppose I say

a 1 is equal to np dot array and l I

create the same array out of this list.

I say print

original

array

right I say A1 right this is my original

array same as this now I want to square

it

square the elements of arrays very very

simple nothing you require you just say

you just say

sq

a1 is equal to sq a1 1 is equal to a1 to

the power of 2, right? A1 to the power

of 2, right? And then you print

the

right you get the squared r. Suppose

you want to find the mean of

the mean of numbers

using list. Right? So what will you do?

You will say

mean is equal to sum of

l right list divided by length of list

right and you will get the means

right which is three for this one right

this original list. Now let me show you

in arrays

right. How will you do this? You will

say mean is equal to np dot mean and you

will just pass a1 right and when you

will check mean you will get 3.2

Right? Nothing like this direct right

direct like this right you can you have

I've already showed you add I've already

showed you uh square and then I believe

you can understand that what all

operations are possible using the arrays

right leveraging the power of arrays

right also guys in the arrays right in

the arrays what you can do is you can

perform

you can perform some string operations

right very powerful string operations

right so for example let's say I have a

I I have a array right I have an array

of names of people right so I say names

is equal to np dot array right and I say

Radha

uh then I say D

right and then I say

Maduk right these three names I have

right so you can check the names they

will be like in the array right and the

data type will be U6 which is a

representation of strings right now guys

suppose you want to capitalize you wish

to capitalize the names right of people

what will you do you will say print me

np docare right npcare dot capitalize

right capitalize and inside this you

will pass names

and you will see all the names have been

capitalized

right you see this R has been

capitalized D has been capitalized. M

has been capitalized.

Right?

You can convert them into upper if you

want. Print

np.care dot upper

names and you will have all of them in

caps lock. You can say print np.care

dot lower.

You will have them in lower.

Right? You can put the title. Right?

Suppose I say uh

title e title is equal to np dot array

and I say

rahov

go

right then I say

dhapati

right and I say mad

warm

right I say these three things now if I

say print

npcare

dot

title right and I say

titles you will see that all the words

will be in capitalized mode rael d and s

of dh sinapati m and V of MaduMa are now

capitalized. I can also

replace something if I wish to suppose I

want to replace Madhu with say suri

right I will say print

right np do.care care dot replace

right and you will say where you want to

replace I say title

and in this I want to replace

madu

with

suri

right and you will say it will be suri_1

suri one Right. Madu has been replaced

with

Right.

Right. If you want to calculate

the characters of strings,

right, you can do that. Print np.car

dot str length of

titles and you will get 11 characters

are there in here.

15 are there here and 11 are here in

this particular thing.

Right? You can do much more powerful

things also. Let me show you one complex

function. Right? Suppose I have f name

is equal to np dot array.

Right? And we have Ra

Madu

right and we have L name

arapati

worma. Right, we have these two things.

Now I can create a new array full name

by simply right by simply saying np.car

car dot add right and I say uh f name

right

comma

l

name right

Yes.

Yep.

Full name.

Yeah. Ra. I just was trying to add a

space

in between.

Anyway,

right. We can do that,

right? You can also people search in

arrays, right? Very powerful. Again,

search in arrays using where,

right?

Right. You can say suppose a 2 is equal

to

np dot array right and I'll say 1 1 22

33 3 4 4 5 66 right and now you have to

say a is equal to np dot where right and

in this you say a2 greater than 20

right and when you will check A

uh it's giving me the index. Why is it

giving me the index

or is it giving me the index?

Does it always return index?

One more thing is you can find a you can

find an

a letter

right through a letter. You can find an

element

through

a letter. Guys, these are all some

tricks which you should know because you

will be dealing with data and you need

to pull data, right? You need to pull

data a lot, right? Based on conditions

and based on things. Suppose you want to

find out the names which have G in them

right or R A in them. So how will you do

this? I will say print right and I will

say uh np do.care care dot find right

and I will say find this inside full

names right and find me ra a right ra a

two p

it uh right it returns true because it

has found it here right so I don't want

to tell you indexing through this but

anyway you should know this just just

assume this that I'm telling you to

write this okay because this is much

easier when we will go to pandas right

just uh write it as a syntax okay

greater than equal to zero I hope this

is clear

>> so let's start with the data science

interview questions and answers and The

number one problem we would be facing is

real world problem solving. And the

question one is handling missing data in

predictive modeling. So imagine you have

given a data set where 30% of the data

for key predictive variable is missing.

This variable is crucial for a

predictive model. How would you handle

this situation to ensure the integrity

and performance of your model? And

please describe your approach step by

step. So starting with the answer you

can start with handling missing data set

is a common challenge in data science

and it's important to address it

carefully to maintain the accuracy of

your model and here's how you could

approach this situation. The number one

point could be identify the missing

data. So first you need to understand

where the missing values are in your

data set. You can do this by using a

simple code in Python with libraries

like mandas. For example, you can use

the data dot isnull dot sum function

that will show you the count of missing

values in each column. Then you can

analyze the pattern. Determine if

there's a pattern to the missing data.

Is it random or is it missing for a

reason? This can affect your approach.

If the data is missing at random, the

methods you use might be different than

if the data is missing systematically.

So choosing a method for imputation.

Let's see the next method that is

choosing a method for imputation. So if

the missing data is numeric, you might

replace missing values with the mean or

median of that column. This is simple

and effective but can be used primarily

when the data is missing completely at

random. Then comes model based

imputation. Sometimes you can use other

variables in the data to predict missing

values using a regression model. This

can be more accurate but is also more

complex. Then we'll use the k nearest

neighbors can algorithm. But before that

we have a code snippet here that could

be used for the implementation of

imputation. You could use Python or R.

And now moving on we'll see the K

nearest neighbors algorithm. So this

method predicts the missing values based

on how closely related the data points

are to each other. So after imputation

it's crucial to check how your changes

have affected the overall data set and

model performance. Sometimes filling in

too many missing values can introduce

bias. And then we have visualization. To

help understand before and after the

imputation, you could visualize the

distribution of the variable using

histograms or box plots. This helps in

seeing how the imputation has changed

the statistical properties of the data.

And by following these steps, you can

handle missing data thoughtfully and

maintain the integrity of your

predictive model. Now moving to the

question number two that is based on

evaluating model overfitting. So the

question is you have developed a

predictive model but you suspect it

might be overfitting the training data.

How would you test and address the

issue? Please explain your steps and the

techniques you would use. So you could

start the answer by explaining what is

overfitting. So overfitting is a common

problem where model performs well on

training data but poorly on unseen data

indicating it's too closely fitted to

the training data specific details and

noise. So now we'll see a step-by-step

guide on how to address this. The number

one step is cross validation. So one

effective way to test for overfitting is

by using cross validation technique.

Cross validation involves splitting your

training data into multiple smaller sets

that is false and then training a model

on some of these set and validating it

on the others. So this helps you

understand if the model's good

performance is consistent across

different subsets of data. For example,

in Python you can use the cross value

score function from skarn.mmodel

selection. So this is the code and this

is the code snippet of Python that you

can use for the cross validation and

here we are importing from skarn that is

the module and we're importing

cross_well

score and here we have used the cross

val score function and then we have

printed the average cross validation

score and the next step we will do is

running cross validation model. So this

is your predictive model that you have

already built using scikit learn and

here's the x train these are the x input

features of your training data and y

train these are the output labels of

training data. So we are running gross

validation model here this is your

predictive model that you have already

built using scikitlearn. So x train here

that means these are the input features

of your training data and y train here

means these are the output labels of

training data and cv equal to 5. This

parameter tests the function to split

the data into five parts that is false.

And the model is trained on four of

these parts and the remaining part is

used for testing. So this process

rotates until each part has been used

for testing once and the printing

results that is score dot mean. So this

calculates the average of the scores

obtained from each gross validation for

this average score gives you an idea of

how well your model is likely to perform

on unseen data. A consistent score

across different polls suggests your

model is generalizing well rather than

overfitting to the training data. So now

moving to the next point that is

training versus validation error. So

plot the training and validation errors

as a function of training epochs or

complexity of the model. A model that

overfits will show a low error on

training data and a high error on

validation data as it trains further.

Then we have pruning the model. If you

confirm that the model is overfitting,

consider simplifying it. This might mean

reducing the number of parameters by

selecting fewer features using

regularization techniques like lasso or

ridge or choosing a less complex model.

After this step, we will move to

regularization technique step. So these

techniques add a penalty to the loss

function used to train the model which

can discourage complex models that

overfeit. Then we have common methods

that include L1 that is lasso and L2

ridge regularization. And here's how you

can add L2 regularization in Python. So

this is the code snippet here. And what

we have done here is we are creating the

ridge model and we have applied alpha

equal to 1.0. So this parameter controls

the strength of the regularization. A

higher alpha value increases the

regularization effect which helps reduce

model complexity and combat overfitting.

The alpha value can be tuned to find the

optimal balance between bias and

variance. And now coming for the fitting

the model. So model do fit and in that

we have X train and Y train that trains

the ridge model on the training data. It

adjusts the weight of the feature in X

train to predict the Y train while also

considering the regularization term.

This helps prevent the model from

fitting too closely to the noisy aspects

of the training data. And then we are

re-evaluating the model. After making

adjustments, it's important to

re-evaluate the model again using the

same cross validation technique to see

if the issue of overfitting has

improved. And then we have

visualization. To help illustrate or

ffitting, you could create a plot

showing the training and validation

errors or the number of epochs or model

complexity. So by using these

techniques, you can identify if your

model is all fitting and take steps to

correct it ensuring it performs well not

only on the training data but also on

new unseen data. So now moving to the

next question that is question number

three and it is based on realtime data

stream processing and the question is

you are tasked with building a model to

predict stock prices in real time. The

data comes in every second and you need

to update your predictions accordingly.

Describe how you would set up your

system to handle this type of data

effectively and what tools and

techniques would you use and why. So you

could start answering this question with

handling real-time data. So handling

real-time data especially for something

as volatile and fastpaced as stock

prices requires a robust system that can

process and analyze data quickly and

accurately. So here's how you could

approach this. We will set up such a

system and we'll have some steps. So

starting with the steps. So the first

step is choosing the right tools. The

right tool would be Apache Kafka. So

this is a popular tool for handling

real-time data that streams because it

allows you to publish and subscribe to

streams of records that is data and it

can handle high throughput with low

latency. Kafka acts as a buffer and

manages the flow of data ensuring that

your system doesn't get overwhelmed and

you can also use Apache Spark especially

Spark streaming is excellent for

processing the data. It can process data

in real time and perform complex

operations like windowing, grouping data

into chunks of a specified time period

and aggregating that is summarizing

data. So you can modify it and perform

the predicting of stock prices. And then

the step is data processing pipeline.

And the first step comes here is

injection. Data first enters the system

typically through Kafka which collects

data sent from the stock market and then

we do the processing. So spark streaming

takes over here. Here you can apply

transformations and run your predictive

models on the data. For example, you

might calculate moving averages or other

indicators that feed into your stock

price prediction model. And then comes

the output. Finally, the predictions are

outputed. This could be to a dashboard

for traders, an automated trading system

or even stored for further analysis. And

then we develop the model. Now comes the

model development. You would likely use

a machine learning model that can update

quickly and incorporate new data as it

arrives. models such as aim for time

series forecasting or more complex

machine learning models like rect neural

networks RNNs can be suitable. The model

should be retrained or fine-tuned

periodically with new data to ensure it

stays accurate. Now we'll come to

scalability and reliability. So ensure

your system can scale as data volume

increases. This might mean adding more

servers or optimizing your data

processing code. Implement monitoring to

catch any issues early like delays in

data processing or model performance

drops. And now we'll see the step that

is visualization and monitoring.

Consider setting up a real-time

dashboard that shows key metrics like

prediction accuracy and processing time.

This helps in quickly spotting when

something goes wrong. By setting up your

system with these tools and strategies,

you can effectively handle the challenge

of predicting stock prices in real time.

So now we'll move to the next question

that is question number four and this

will based on scalable data analytics.

So we have covered two questions that

were a bit code based questions and now

we'll see other questions that would be

based on scalable data analytics or they

might be on different areas and with the

13th question we'll start again with the

coding ones. So moving with the question

four that is based on scalable data

analytics and the question is given a

scenario where your organization

suddenly needs to scale its data

analysis capabilities due to an influx

of data that would be 10 times the

normal volume. How would you handle this

situation to ensure your data analytics

processes remain efficient and accurate?

What technologies would you consider and

what steps would you take? So you can

start answering this question with

handling a sudden increase in data

volume requires a strategic approach to

scaling your analytics infrastructure

without compromising on efficiency or

accuracy. So we'll see some steps from

that you could effectively manage this

scenario that you would start answering

the interviewer that we can start by

evaluating the current infrastructure's

ability to handle increased loads. This

includes assessing your databases,

servers and analytical tools to identify

potential bottlenecks or limitations.

Then you could move to next step that

would be choosing scalable technologies

to manage the increased data volume.

Consider leveraging cloud-based

solutions such as Amazon web services,

Google cloud platform or Microsoft

Azure. These platforms offer scalable

resources which can be adjusted

accordingly to the data load ensuring

you only pay for what you use. integrate

big data technologies like Apache Hadoop

for distributed storage and Apache Spark

for fast data processing. These tools

are designed to handle massive volumes

of data efficiently and can scale up to

meet standard increased demands. Now we

move to the next step that would be

optimizing data processing. So implement

data partitioning and indexing

strategies to improve the efficiency of

data queries. This will help in managing

large data sets by breaking them into

smaller manageable chunks and speeding

up search operations and use real-time

data processing frameworks like Apache

Kafka or Apache Flink which can handle

high throughput and provide timely

insights from large data streams. And

the next step would be automation and

monitoring. Automate routine data

processing task to reduce the manual

effort and speed up the analysis. This

can be done through scripting or using

workflow automation tools. Set up

comprehensive monitoring systems to

track the performance of your data

processes. Tools like Prometheus for

system monitoring and Graphana for

analytics and monitoring dashboards are

useful here. They help ensure that the

system is running smoothly and alert you

to potential issues before they become

critical. And the next step will be

regular evaluation and scaling.

Continuously evaluate the performance of

analytics infrastructure. As your data

grows, keep adjusting and scaling your

resources to maintain optimal

performance. Plan for periodic reviews

of your technology stack and

infrastructure to ensure they remain

aligned with your data needs and

organizational goals. By following these

steps, you can ensure that your data

analytics processes are prepared to

handle sudden surges in data volume

effectively maintaining the integrity

and speed of insights. So this was all

for the question four. Now moving to the

question five and this is based on

integrating machine learning models into

production and the question is you have

developed a machine learning model that

performs well in testing environment.

Now you need to integrate it into your

production environment where it will be

used in realtime applications. What

steps would you take to ensure the

successful deployment and operations of

the model in production? So we'll start

answering this by successfully deploying

a machine learning model into production

involves several critical steps to

ensure it performs as well in real time

operations as it does in testing. So you

would have a clear pathway to make the

interviewer understand. We will start

with the pathway with the first step

that would be model validation. So

before moving anything into production

revalidate your model's performance

using a separate validation data set.

This helps confirm that the model

generalizes well to new unseen data. The

next step will be preparing the

production environment. Ensure that the

production environment is ready to

handle the model. This includes setting

up the necessary hardware and software

ensuring that it can handle the expected

load and that all dependencies are

correctly installed and configured. Then

the next step comes that is model

wrapping. Wrap your model in an API that

is application programming interface

making it accessible to other parts of

your software infrastructure. Frameworks

like flask for Python can be used to

create a simple web server that listens

for data inputs and provides model

outputs. Then comes the next step that

is deployment strategies. Consider using

containerization tools like doer which

can help encapsulate your model and its

environment ensuring that it works

uniformly across different development

and production settings. And then we'll

use deployment strategies like blue

green deployment or canary releases to

minimize downtime and reduce the risk of

introducing a faulty model into

production. And then comes the next step

that is monitoring and logging.

Implement logging and monitoring to

track the model's performance and health

in real time. Tools like prompts for

monitoring and ELK elastic search log

statch kibbana for logging help in

quickly identifying and diagnosing

issues in production. And then comes the

next step that is performance tuning.

Monitor the model's performance over

time. If the model's performance

degrades or if new data shows different

patterns, you may need to retrain or

fine-tune the model to maintain

accuracy. And after this step, there's a

step for feedback loop. Set a feedback

loop where predictions and outcomes can

be compared. This feedback is crucial

for continuously improving the model and

catching any drift in data or changes in

external conditions that affect the

model. And after this comes a last step

that is legal and compliance checks.

Ensure all the data used by the model in

production complies with privacy laws

and regulations. This is crucial for

maintaining trust and legality

especially when handling sensitive

information. So by carefully planning

and executing these steps you can

smoothly transition your machine

learning model from a testing

environment to a fully functional

component of a production system. So

this was all about the question number

five. Now moving to the question number

six that would be based on datadriven

decision making. And the question is

your company wants to shift towards more

datadriven decision making. You have

been tasked with developing a strategy

to implement this. What steps would you

take to ensure that the data at all

levels of the organization is utilized

effectively to make informed decisions

and what challenges might you face and

how would you address them? So you can

start answering this by implementing a

datadriven decision-m strategy that will

require a comprehensive approach to

ensure that reliable data is accessible

and effectively used across all levels

of the organization. And now we can

develop and deploy this strategy. And

similarly you could tell this strategy

to the interviewer. So the number one

step will be assessing current data

infrastructure. Start by evaluating the

existing data infrastructure to

understand what data is available, how

it is stored and how it is currently

used. This assessment will help identify

gaps in data collection, storage and

access that need to be addressed. Now we

move to the next step that is developing

a data governance framework. Implement a

data governance framework that defines

who can access data, how it can be used

and who is responsible for maintaining

its quality. This framework ensures data

integrity and security which are

critical for making reliable decisions.

Now we'll move to the next step that is

training and empowerment. So train

employees at all levels on the

importance of datadriven decision making

and provide them with the tools and

knowledge necessary to analyze and

interpret data. This might include

training sessions, workshops and ongoing

support to ensure everyone can use data

effectively. Now move to the next step

that is implementing analytical tools.

So deploy user-friendly analytical tools

that can integrate seamlessly into the

daily workflows of employees. Tools like

Tableau, Microsoft PowerBI or even

advanced Excel techniques can provide

powerful data analysis capabilities

without requiring extensive technical

knowledge. After this we'll move to the

step that would be creating a

centralized data platform. Developer

centralized data platform where all

organizational data can be accessed and

analyzed. This platform should be

scalable and secure providing a single

source of truth for the organization.

And then we have the promoting a

datadriven culture. So foster culture

that values datadriven decision-m

encourage experimentation and learning

from datadriven initiatives. celebrate

successes and learn from failures to

continually improve the use of

datadriven in decision making and there

would be some challenges and solutions

for that. So one major challenge we know

here is resistance to change as some

employees may prefer traditional

decision-m methods. So address this by

demonstrating the tangible benefits of

datadriven decisions through pilot

projects and success stories. So data

silos can also hinder effective data use

promote cross department collaboration

and integrate disparate data sources to

overcome this challenge. After that you

can monitor and do continuous

improvement. So by systematically

implementing these steps you can

transform your organization into one

that leverages data at all levels to

make informed and effective decisions.

And after answering in these steps you

could make the interviewer have a truth

and a faith in you that you could make

these models. Now move to the next

question that is question number seven

and that is based on handling large data

set and the question is your project

involves analyzing extremely large data

sets potentially exceeding terabytes in

size. What strategies would you use to

manage and analyze such large data sets

effectively? Describe the tools and

techniques you might employ and you

could start this with answering that

working with large data sets especially

those in terabyte range presents unique

challenges in terms of storage

processing and analysis. So we'll have a

structured approach to handle these

challenges effectively. We'll start with

the data storage that would be use

distributed file systems. Consider using

a distributed file systems like Hadoop

distributed file system HDFS or Amazon

S3. These systems are designed to store

vast amounts of data across many servers

offering high availability and port

tolerance. And then comes the next step

that is data processing. Leverage big

data processing frameworks. Tools like

Apache Spark are ideal for processing

large data sets because they handle

distributed computing effectively. Spark

can perform data processing task much

faster than traditional disk based

processing due to its in-memory

computing capabilities. And next we

could start with efficient data

sampling. So there are many sampling

techniques that we can use. So when the

data set is too large to handle even

with powerful tools consider using data

sampling techniques to reduce the size

to a manageable level without losing

significant insights. Ensure that the

sample represents the whole data set

accurately. And then comes optimization

of data queries. Indexing and

partitioning. Optimize your data queries

by implementing indexing and

partitioning. This can drastically

reduce the time it takes to perform

queries by limiting the amounts of data

scan. And then we can do scalable

analytics. And then we'll move to the

next step that is scalable analytics.

And in that we could start with the

parallel computing. Use parallel

computing capabilities of frameworks

like spark or dask to analyze data

across multiple nodes. This helps in

scaling up your analytics operations to

handle large data sets effectively. And

now we'll move to the cloud-based

analytical tools. So consider using

cloud services like Google BigQuery or

AWS Red Shift which are designed to

handle massive data sets and complex

analytics with ease. And after this step

we'll move to data cleaning and

pre-processing. Here we will automate

pre-processing task. We'll use automated

tools to clean and pre-process data.

This includes handling missing values,

normalizing data and removing duplicates

which can be particularly challenging

with large data set. And after this

step, we'll move to the step that will

visualize large data set. So we'll use

specialized tools. That tools could be

Tableau or PowerBI that can handle large

data set by aggregating data and using

efficient backend technologies. For more

detailed exploration, tools like plotly

or bouquet can be used as they offer

capabilities to interactively visualize

large volumes of data. And after that,

there would be step for regular

maintenance and updates. That could be

continuously monitoring the data

quality. As new data comes in, you can

continuously monitor its quality. And

after this step, you could integrate all

these strategies and tools into your

workflow. And you can effectively manage

and extract valuable insights from

extremely large data sets thereby

supporting robust datadriven decision

making. And you could answer the whole

strategy to the interviewer. Now moving

to the question number eight that is

based on optimizing machine learning

models and the question is during model

development you have noticed that your

machine learning model is

underperforming. What steps would you

take to diagnose the problem and

optimize the model's performance? What

techniques and tools would you use? So

you can answer this by starting with the

optimizing and optimizing a machine

learning model that is underperforming

involves several steps to diagnose and

improve its accuracy and efficiency. And

here we will have structured approach to

tackle this issue and you could start

this with the number one step that is

diagnosing the problem. Evaluate model

metrics. Start by thoroughly evaluating

the performance metrics of your model.

For classification task, for

classification task, look at accuracy,

precision, recall and the F1 score. For

regression task, consider R squ, mean

squared error that is MSE and mean

absolute error that is MA. And then you

can move to the next step that is use

plots like ROC curves for classification

models and residual plots for regression

to visually assess with the model is

going wrong. After that, we'll move to

the next step that is data quality and

quantity check. Inspect the data that is

sometimes the quality and quantity of

data can be the root cause of poor model

performance. Ensure the data is clean,

well pre-processed and sufficient. Look

for issues like missing values, outliers

or imbalanced classes. And after this

we'll move to the feature engineering

step that would be experiment with

creating new features or transforming

existing ones to provide better

predictive power. And then we have the

next step that is model tuning and

configuration. After feature

engineering, we'll move to the next step

that is model tuning and configuration.

So, hyperparameter tuning. Use

techniques like grid search or random

search to find the optimal settings for

your model's parameters. Tools like

scikit learns, grid search CV or

randomized search CV can automate this

process. And there's a cross validation

that would implement cross validation to

ensure that the model's performance is

consistent across different subsets of

the data set. And then we have the next

step that is trying different models. So

experiment with algorithms here. If

initial models are underperforming, try

different algorithms that might be

better suited for the problem. For

instance, if you started with linear

regression and it's not performing well,

consider more complex models like random

forest or gradient boosting machines.

And after this, we have nseml methods

that we can use techniques like bagging,

boosting or stacking to combine the

predictions of multiple models to

improve overall performance. After this

step, we have feature selection that

includes reduce dimensionality. Use

techniques like principal component

analysis that is PCA to reduce the

number of features which might help in

improving model performance by removing

noise and redundancy. And then we have

select important features. So use model

based technique to identify and keep

only the most important features that

impact the outcome. And then comes the

last step that is regular updates and

retraining. So here you can monitor and

update that could be continuously

monitoring the model's performance over

time as new data becomes available

update and retrain the model to adapt to

any changes in underlying patterns and

after that you could have a consultation

and collaboration work with the other

teams and by methodically addressing

each of these areas you can diagnose why

your machine learning model is

underperforming and can take steps to

optimize its accuracy and efficiency. So

this was all about question number

eight. So let's start with the question

number nine and this is based on

handling unstructured data. So the

question is you are given a large amount

of unstructured data including text,

images and videos. What strategies would

you use to manage and analyze this type

of data effectively? Describe the tools

and techniques you might employ. So you

can start answering this question by

describing that dealing with

unstructured data can be challenging due

to its lack of predefined format or

structure. However, with the right

strategies and tools, you can

effectively manage and analyze it to

extract valuable insights and there will

be a approach how you can do that. So,

we will discuss the approach here and

starting with the steps. So, the number

one step will be data categorization and

organization. So, the number one step in

this step will be sorting and tagging.

We will begin by categorizing the data

into types that will be text, images or

videos. Use tagging to add metadata

which helps in organizing the data and

makes it easier to access and analyze

later. Then and after that particularly

for text data we'll use natural language

processing NLP. We will employ NLP

techniques to extract useful information

from text. Tools like NLTK, spacy or

even more advanced models like BERT can

help you perform tasks such as sentiment

analysis, entity recognition and topic

modeling. After that we will do text

indexing. We can use elastic search or

Apache sle to index large volumes of

text. These tools provide powerful

search capabilities and can handle

complex queries efficiently. And after

that we'll move to image data. And to

structure image data we'll use image

processing. We'll use libraries like

OpenCV for basic image processing tasks

such as filtering and transformations.

For more advanced image analysis,

consider deep learning models using

frameworks like TensorFlow or PyTorch.

And then we'll feature extraction. Apply

techniques to extract features from

images such as edges, textures or key

points which can be used for further

analysis or machine learning. And then

we'll come to video data. And here we'll

do video processing. and we'll use the

tools like fmpg that can be used for

basic video processing tasks such as

format conversion or extracting frames

for analyzing video content look at

machine learning models that can

classify or recognize activities in the

video and after this we'll move to

temporal analysis for videos temporal

components are important techniques like

sequence modeling or recurrent neural

networks RNNs can be useful to analyze

sequences of frames for activities or

events and then we'll move to data

storage and management. Here we'll use

the given volume and complexity of

unstructured data and use big data

platforms like Hadoop or cloud services

like AWS S3 for storage. These platforms

can scale up to handle large data sizes

and provide the necessary infrastructure

to store and retrieve unstructured data

efficiently. And then we have

visualization and reporting custom

dashboards that we'll create here. We

will develop custom dashboards using

tools like Tableau or PowerBI which can

integrate different data types and

provide a unified view of the analyzed

data. And after that we will do data

summarization. Tools that provide

summarization capabilities can help in

considering large volumes of

unstructured data into more manageable

and interpretable forms. And after that

we'll leverage these strategies and

tools and can effectively manage,

analyze and derive insights from

unstructured data which can be crucial

for making informed decisions in various

applications. And this is the path that

you can explore and explain to the

interviewer if this question has been

asked. Now moving to the question number

10 and that will be based on scaling AI

solutions in enterprise and the question

is your company wants to scale its AI

operations from a few initial pilot

projects to enterprisewide

implementation. What are the key

considerations and steps you would take

to ensure the successful scaling of AI

solutions across the organization and

what challenges might you face and how

would you address them? So you can start

answering this question with the scaling

AI solutions. You could answer him that

scaling AI solutions across an

enterprise requires careful planning and

strategic implementation to ensure

success and alignment with business

objectives and there should be a

strategic approach to implement this. So

starting with the approach and the

number one step will be that will be

strategic alignment. So identify

business objectives. Start by

identifying the business objectives that

the AI solutions are intended to

support. This ensures that the AI

initiatives are aligned with the company

strategic goals and can demonstrate

clear business value. And then comes the

stakeholder engagement. So engage

stakeholders from various departments

early in the process to gather input and

build support. This helps in

understanding diverse needs and ensures

broader acceptance of the AI solutions.

And after that comes the infrastructure

and technology. So there's an option

that is assess and upgrade

infrastructure. Evaluate whether your

current IT infrastructure can support

the expanded use of AI. You might need

to upgrade hardware, invest in cloud

solutions or adopt technologies that

facilitate AI processing and data

handling. And after that we have

standardization of tools. Standardize

the tools and platforms used for AI

development to ensure compatibility and

ease of maintenance across the

organizations. And after that we'll move

to data management. So robust data

governance that is to implement a strong

data governance framework to manage

enterprise data effectively. This

includes policies for data quality,

security and compliance especially

important when scaling AI solutions that

rely on vast amounts of data. And after

that we will come to data accessibility.

So ensure that data is accessible across

the organization but also secure against

unauthorized access. This involves

setting up secure data leaks or

warehouses that centralize data while

allowing controlled access. And then we

come to the next step that is talent and

training. So build AI competency that is

develop in-house AI expertise through

training programs and hiring. So this

build the necessary skills within the

organization to develop, manage and

scale AI solutions and after that you

can also perform cross functional AI

teams that could be forming cross

functional teams that include data

scientists, IT professionals and domain

experts. So this fosters collaboration

and ensure that AI solutions are

developed with a comprehensive

understanding. And after forming these

collaborative teams, we move to scalable

deployment models. So pilot test and

phase roll out. Before a full-scale

rollout, conduct pilot test to go the AI

solution effectiveness and integration

capabilities based on feedback, adjust

and then gradually deploy the solutions

across the organization. And then we

have modular and flexible design. So

design AI systems to be modular and

scalable allowing for adjustments and

expansions as needs and then we'll

monitor and do the continuous

improvement. So there will be

performance metrics that would establish

metrics to regularly assess the

performance of AI systems. We will

monitor these systems to ensure they met

expected outcomes and adapt as

necessary. And after that we have next

step that is addressing challenges. So

there could be cultural resistance that

there could be employees that would be

resisting to the changes but we have to

address this through continuous

education and by showcasing successful

AI use cases within the organizations

and by carefully considering these

aspects and methodically implementing

steps you can successfully scale AI

solutions across your enterprise driving

significant business value and

innovation. And that's all for question

number 10. Now we'll move to question

number 11 and that is based on ethical

considerations in data science. So the

question is in your data science

projects how do you ensure that ethical

considerations are addressed? Describe

the steps you take to identify and

mitigate ethical risk in your projects.

What frameworks or guidelines do you

follow? So you could start answering

this question with ethical

considerations that they're crucial in

data science to ensure that the

solutions and analyzes do not

advertently cause harm or bias. Here's

how you can ensure that. So there are

some steps and we will discuss those

steps. Starting with the number one that

is educate on ethical standards. So stay

informed about the ethical standards in

data science such as fairness,

accountability, transparency and

privacy. Organizations like the data

science association and the ACM have

codes of ethics that we refer to as

guidelines. And then we have ethical

risk assessment. Identify potential

ethical issues. That would be at the

beginning of each project. Conduct a

thorough assessment to identify any

potential ethical risk such as biases in

data or impact on vulnerable groups.

This involve reviewing the source of

data, the methodologies used for data

collection and the intended use of the

data analytics results. And then we have

stakeholder analysis. Engage with

stakeholders to understand the diverse

perspectives and potential impact of the

project. This helps in identifying

ethical issues that may not be apparent

from a purely technical standpoint. And

then we'll move to mitigation

strategies. Implementing bias mitigation

techniques. We will use statistical and

machine learning techniques to detect

and mitigate biases in data. This might

involve techniques like resampling,

reeing or using algorithms designed to

be fair. And then we have privacy

preserving methods. Employ methods such

as data anonymization, encryption or

differential privacy to protect

individual privacy when analyzing

sensitive data. Then we have other

methods that is transparency and

explanability. There we have model

explanability and after that coming to

documentation and reporting. So we have

to maintain thorough documentation of

data sources, model decisions and

methodologies. And then we have

continuous monitoring and feedback.

There you have to monitor outcomes and

the feedback mechanisms should be

applied. And then we have the panels

that is collaboration and advisory

panels. Then we have ethical review

boards. So for complex projects setting

up or consulting with an ethical review

board can provide oversight and diverse

perspectives on the ethical implications

of project methodologies. So by

proactively addressing ethical

considerations through these steps you

can ensure that your data science

projects uphold high ethical standards

and positively contribute to society

while minimizing harm. So this was all

about question 11. Now moving to

question number 12 that is based on time

series forecasting for business

decisions. So the question number 12 is

you are tasked with forecasting monthly

sales for a retail company using time

series data from the past 5 years. What

steps would you take to prepare and

analyze this data to make accurate

forecast? What specific tools or

techniques would you use and why? So we

can start answering this by time series

forecasting and we could address them

that it's a powerful tool for predicting

future events based on past data

especially in business context like

retail sales. So we will have a

structured approach here and we'll start

with data collection and cleaning.

First, you will gather data and ensure

that you have collected all relevant

data including monthly sale figures from

the past five years and also considering

including external factors that might

affect sales such as economic

indicators, holidays and promotional

activities. And then we'll proceed to

clean data. We will check for and handle

any inconsistencies or missing values.

And then we have data visualization.

Here we will plot the data. We'll use

plotting libraries like Matt Lib or

Seabbone in Python to visualize the

data. This will help in identifying

patterns, trends and seasonality. And

then we have decomposition of data. So

there's a seasonal decomposition and

we'll use statistical techniques to

decompose the data into trend seasonally

and residuals. So this can be

accomplished with tools like the

seasonal decompose function from the

stats models library in Python. And

we'll understand these components

separately and can improve the accuracy

of our forecast. And then the next step

is model selection and forecasting. So

there are two models that is a ara and s

IMA models. So we have to choose

appropriate forecasting models based on

data's characteristics. For instance,

AMA that is auto reggressive integrated

moving average. It is effective for

non-season data while SMA that is

seasonal AMA that is suitable for data

with seasonal patterns. And after

choosing the model we'll move to cross

validation. We will implement time

series specific cross validation

techniques like timebased splitting to

evaluate model performance and this will

ensure your model generalizes well on

unseen data and then we have model

fitting and diagnostics. We will fit the

model that is by using the cinemax class

from stat models that will fit your

model to the data and then we will

carefully select parameters based on AIC

that is a cake information criterion

that scores or thorough grid research

technique and then we can do the

diagnostics and forecast and validation

and after forecast validation we'll move

to iterative improvement. So there's a

feedback loop that should be mandatory

and there should be a regular update for

the model with new sales data and

refining the model as needed. So this

continuous improvement cycle helps adapt

to changing patterns in sales data. And

by following these steps and using these

tools, you can create robust forecast

that help the retail company plan better

and make informed decisions. So this was

all about question number 12. Now move

to question number 13 that is based on

customer segmentation using machine

learning. So the question is you are

given a data set containing demographic

and purchasing behavior data for a group

of customers. Your task is to segment

these customers into distinct groups

based on similarities in the purchasing

behavior and demographics. So what steps

would you take to perform this

segmentation and can you provide a

sample Python code snippet to illustrate

the initial stages of data handling and

model application. So we can start this

by explaining customer segmentation that

it's a powerful approach to tailor

marketing strategies and improve

customer service by identifying distinct

groups based on their behavior and

characteristics. And here also we have a

detailed approach for this task. So

we'll start with number one step that

would be data exploration and

pre-processing. So there will be initial

exploration that is beginning by

examining the data set to understand the

features available such as age, income,

purchase frequency etc. Then we'll look

for missing values or anomalies and

decide how to handle them. That could be

using imputation and then we'll move to

feature engineering. We'll create new

features that might be useful for

segmentation such as customer lifetime

value or average transaction amount.

We'll also use normalization that is

normalize the data to ensure that one

feature doesn't disproportionately

influence the model due to its scale.

We'll use standard scaling or minmax

scaling as appropriate. So then we'll

come to the next step that is choosing

the segmentation technique. And here we

have k means clustering. So this is a

popular method for customer

segmentation. Here we will decide on the

number of clusters by using techniques

like the elbow method or analysis to

determine the optimal cluster count. And

then we have model implementation and in

that we will use data preparation and

we'll prepare the data by selecting the

relevant features and applying any final

transformations and then we have model

fitting. We fit the C means clustering

model to the data and evaluate and

interpret analyzing clusters and after

analyzing clusters we'll move to the

next step that is strategic insights. We

will provide actionable insights based

on cluster characteristics such as

targeted marketing strategies for each

segment. And then we have iterative

refinement that is feedback

incorporation and we'll use business

feedback to refine the segmentation. If

additional data becomes available

incorporated to enhance the model and

now we'll see the sample Python code. So

for this first we'll import the

libraries and modules. As you can see on

the screen we have imported pandas

random forest classifier train test

split standard scaler classification

report and after that we will load the

data and and for that we have used the

pandas to read the data that is read ssv

and after that we are processing the

data that is data prep-processing we are

handling missing values and using the

forward fill or fil to fill missing

values in the data set and then we are

featuring scaling that is normalizing

the selected features that is feature

one, feature two and feature three using

standard scaler and then we'll move to

the next step that is data splitting.

We'll split the data set into training

and testing sets. So that test size

equal to 0.2 parameters specifies that

20% of the data will be used for testing

and then we'll train the model. We'll

initialize and train a random forest

classifier with 100 trees and a random

state for reproductibility and then

we'll evaluate the model. will make

predictions on the test set that is X

test using the train model and print a

classification report showing precision

recall F1 score and support for each

class. So this code demonstrates the

process of loading, pre-processing,

training and evaluating a machine

learning model that is random forest

classifier for predicting equipment

failures in a manufacturing plant. The

use of techniques such as data

prep-processing and splitting along with

the random forest classifier highlights

a standard flow for building predictive

maintenance models. So this was all

about the question number 13. So now

moving to the question number 14 that is

based on predictive customer churn and

the question is you are tasked with

developing a model to predict which

customers are likely to churn from a

subscription service. So what steps

would you take to build this model and

can you provide a sample Python code to

illustrate the data preparation and

model training process? So we'll start

answering this question about depicting

what is predicting customer churn. So

predicting customer churn is crucial for

businesses to implement detention

strategies proactively and we'll have a

detailed approach for building a

predictive model for this purpose.

Starting with data collection and

exploration and in this we will collect

data and after that we'll perform the

exploratory data analysis that is EDA.

We'll perform an initial analysis to

understand patterns and trends and then

we have feature engineering. We will

create new features and derive new

feature that might influence churn such

as change in usage pattern or service

upgrades. And then we'll handle the

missing values if we found any. And then

we'll encode categoral variables. We'll

use techniques like one hot encoding or

label encoding for categorial variables.

And then we have scale features to

normalize or standardize numerical

features to ensure they contribute

equally to the model's performance. And

then we'll select the model that is

we'll choose the appropriate model and

start with for the knowing handling

binary classification task that could be

with logistic regression, random forest

or gradient boosting machines. And after

selecting the model, we'll train the

model and evaluate it. So fit your model

on the training data and after that

evaluate the model using appropriate

metrics like accuracy, precision,

recall, F1 score and ROC to go its

performance. And then we'll optimize the

model using hyperparameter tuning. We'll

optimize the model parameter using grid

search or random search to improve

performance. And then we have feature

importance that is analyze and rank

features by their importance in

predicting churn to refine the model

further. And then and then the last step

is deployment and monitoring. We'll

deploy the model once validated deploy

the model into a production environment

where it can predict real-time churn. So

after deploying the model regularly

monitor the model to ensure it remains

effective over time as new data comes

in. So now we'll see the sample Python

code for this example. So starting with

the importing of libraries we will

import pandas numpy scikitlearn skarn

tensorflow and the tensorflow kas and

callbacks and after importing the

modules we'll start with data loading

we'll load the data set from a CSV file

named equipment data dot csv and that

with the pandas data frame and after

that we'll do the data prep-processing

we'll handle missing values and for that

we'll use forward fill to fill missing

values in the data set and then we have

feature scaling that will normalize the

selected features that is feature one,

feature two, feature three using

standard scaler and after that we'll use

the data splitting. We'll split the data

set into training and testing sets and

the test size will be equal to 0.2 and

this parameter specifies that 20% of the

data will be used for testing and after

that we'll start with building the

model. First we'll see sequential model

that initializes a sequential model

technique. And then we have dense layers

that adds two dense layers with 64 units

and value activation function. Then we

have dropout layers that adds two

dropout layers with a dropout rate of

0.5 to reduce overfitting. After that

we'll do the model compilation. We'll

compile the model using the atom

optimizer and binary cross entropy loss

function for binary classification. And

there will be an early stopping that

will define an early stopping call back

to stop training when the validation

loss metric has stopped improving after

three blocks. And after training the

model, we will evaluate the model. And

evaluating the model on the test data

and print the loss and accuracy metrics.

So this code demonstrates the process of

loading, pre-processing, building,

compiling, training and evaluating a

deep learning model using TensorFlow and

KAS for predicting equipment failures in

a manufacturing plant. So the use of

techniques such as data prep-processing,

dropout regularization and early

stopping helps in building a robust deep

learning model for predictive

maintenance. So that's all with question

number 14. Now we'll start with question

number 15 that is based on deep learning

and NLP. And your question is you are

tasked with developing a sentiment

analysis model using deep learning to

understand customer opinions from

reviews. So what steps would you take to

build this model and can you provide a

sample Python code snippet to illustrate

how you would pre-process data and train

a simple deep learning model? So we'll

start answering this with sentiment

analysis that sentiment analysis using

deep learning allows businesses to coach

customer sentiment from text data like

reviews or comments effectively. And

we'll have a detailed approach for

building a sentiment analysis model.

We'll start with data collection and

cleaning. We will collect the data,

gather a substantial data set of text

reviews and their associated sentiments

typically labeled as positive, negative,

or neural. And then we'll clean the

data, pre-process the data by removing

noise such as HTML tags, special

characters, and so words. And we'll

normalize the text by converting it to

lower case. And then we have text

prep-processing. We'll convert text into

tokens, words, or phrases. And then we

have vectorization that transforms

tokens into numerical format using

techniques like word embeddings or TF

that is term frequency in document

frequency and then we'll use the padding

and then we have the option of model

selection. will choose a model

architecture based on a basic approach

and use a RNN or more advanced

architecture like LSTM that is long

short-term memory or GRU that is gated

recurrent units which are effective for

sequence data like text and then we have

model training we'll compile the model

define the model architecture and

compile it with a loss function suited

for classification like categoral cross

entropy and an optimizer like Adam and

then we'll train the model. We'll fit

the model on our pre-processed data.

We'll evaluate and optimize it.

Evaluating model performance. Here use

the metrics such as accuracy, precision,

recall, and F1 score to assess the

model. And then we have hyperparameter

tuning. We'll optimize the model by

adjusting parameters like learning rate,

number of layers and units per layer.

And then coming to deployment. We'll

deploy the model and integrate the model

into the existing review processing

pipeline. So it can automatically

classify new reviews. So let's see the

sample Python code and we'll have a

basic approach for that. Here we'll

import numpy tensorflow sequential

embedding LSTM dense stroke out. So

embedding converts positive integers

that is indexes into dense vectors of

fixed size and LSTM that is long

short-term memory layer that is used for

learning dependencies in sequence data.

And then we have dense that is a

regularly densed connected NN layer. And

then we will import pad sequences. And

after that we have the data set and the

sample text data representing customer

reviews that will store in variable

text. And then we have labels that has

binary labels indicating sentiment one

for positive zero for negative. And now

we'll start with the pre-processing of

data. Here we have declared that

tokenizer. We will initialize a

tokenizer that will help only the top

thousand most frequent words. And then

we have fit_on

text that is update the internal

vocabulary based on the list of text. It

essentially creates a dictionary of word

to index pairs. And then we have text to

sequences that will transform each text

in text to a sequence of integers. And

then we have pad sequences that will

ensure all sequences have the same

length by padding shorter sequences with

zeros up to the maximum length. And then

we'll start building the model. Here we

have sequential model that will set up a

linear stack of layers. And then we have

embedding layer that will map each word

index to an embedding vector of size 64.

So the input length is set to 10 that is

the length of the input sequences. Then

we'll start with LSTM layers. So two

LSTM layers are added. The first one

returns sequences to allow the next LSTM

layer to process these sequences. And

after that we have the dropout layer

that applies dropout with a rate of 0.5

of the first LSTM layer to reduce

overfitting. And after that we'll come

to dense layer that has output of a

single scalar that represents the

predicted setment and using sigmoid

activation to output a probability. And

now we'll start with model compilation

and training. So we'll configure the

model for training and we'll use binary

cross entropy as the loss function that

is suitable for binary classification

and the atom optimizer and tracks

constantly accuracy as a metric and then

we have the fit that trains the model

for a specified number of epochs that is

iterations or the entire data set and

then we'll predict the model that is

after training the model can predict the

sentiment of the reviews in the data

set. This is useful for checking how the

model performs on the training data

itself. So this breakdown explains each

step of the coding process detailing how

the data is prepared and how the model

is configured and then we'll compile it

and use for training and prediction. So

it's detailed explanation should help in

understanding how to implement a simple

LSTM model for sentiment analysis in

TensorFlow. Now moving to the question

number 16. So let's start with question

number 16 that is based on anomly

detection in transaction data. So the

question is you are tasked with

identifying unusual transactions in a

company's financial data that might

suggest fraudulent activity. So what

steps would you take to develop an

anomaly detection model and can you

provide a sample Python code snippet to

illustrate how you would pre-process the

data and apply an anomaly detection

technique. So we'll start answering this

with anomaly detection technique that is

anomaly detection is essential for

preventing fraud by identifying

transactions that deviate significantly

from typical patterns. And now we'll see

the structured approach to building an

anomaly detection model for transaction

data. We'll start with data collection

and cleaning and we'll collect all the

compiling transaction data which should

include details like transaction amount,

time, user ID and transaction type. Then

we'll move to feature engineering and

develop features that capture the

essence of transaction such as time of

day and the day of the week. And then we

have data normalization. We'll use

scaling techniques such as minmax

scaling or standardization to ensure

that the model is perfectly normalized.

And then we have choosing the anomaly

detection technique. So here we have to

choose the technique which is effective

for highdimensional data sets and works

for isolating anomalies instead of

profiling normal data points. After

choosing the anomaly technique will

train anomaly identification. We'll fit

the chosen model to the data and the

anomalies that would have been chosen

will be those transactions that the

model identifies and after this we come

to the last step that is review and

action. Here we have manual review that

is transactions flagged as potential

anomalies should be reviewed manually to

confirm fraudent activity and then we

have continuous improvement that is we

can regularly update the model with the

new data and feedback from the review

process to improve accuracy. And now

moving to the prediction that is after

training the model we can predict the

sentiment of the reviews in the data set

and this is useful for checking how the

model performs on the training data

itself. Now we'll see the Python code to

see how you can set up this model for

anomaly detection. We'll start by

importing the libraries and module and

after that we'll load and prepare data.

That is we'll load transaction data from

a CSV file into the pandas data frame.

And after that we'll convert the

transaction time column to date time

format which allows the extraction of

additional time based features. And

after that we'll perform feature

engineering that will extract the hour

of the day from the transaction time

column. This feature can be important as

transactions occurring at unusual hours

may be indicative of fraud. And then

we'll move to the normalization of data.

This will apply standard scaling to the

amount n of the day feature. This

normalization process involves

subtracting the mean and dividing by the

standard deviation for each feature

ensuring that the feature contribute

equally to the analysis and improving

the performance of many machine learning

algorithms. And after that we'll start

with anomaly detection with isolation

forest. That's a technique. We'll

initialize an isolation forest model

with 100 trees that is n estimators

equal to 100. Setting the proportions of

outliers that is contamination to 1% of

the data. So this parameter is crucial

as it influences the threshold of

marking an observation as an anomaly.

Then we'll fit the model to the scaled

amount n of the data and predict the

anomaly status for each transaction. And

then we'll start with filter and display

anomalies. We'll filter out transactions

identified as anomalies that is anomaly

equal equal to minus one. We'll display

these transactions which can be reviewed

manually to determine if they represent

actual fraud net activity. So this code

snippet provides a systematic approach

to detecting anomalies in transaction

data leveraging the isolation forest

algorithms ability to handle complex and

highdimensional data set effectively. So

the pre-processing steps ensured that

the data is appropriately formatted and

normalized for optional model

performance. So this was all about

question number 16. Now moving to

question number 17 and that is based on

integrating machine learning models into

web applications. And your question is,

you have developed a machine learning

model to predict real estate prices

based on various features like location,

size, and amenities. How would you

integrate this model into a web

application to allow users to get

real-time price predictions? Can you

provide a sample Python code snippet to

illustrate how you would prepare the

model for integration and handle user

request? So, starting with the approach

that is integrating a machine learning

model into a web application. This will

involve several steps to ensure the

model is accessible and perform well in

a live environment. So here's how you

can approach this task. We could divide

into steps and we'll start with number

one step that is model preparation.

We'll finalize and save the model. So

once your model is trained and

validated, save it using a format that

can be easily loaded into a web

application. So Python's pickle module

or TensorFlow's save model format are

commonly used for this purpose. Then we

can use web application backend setup.

For this, select a suitable web

framework. So, Flask is popularly known

for its simplicity and effectiveness in

integrating Python based machine

learning models. And after that, we'll

develop the API. After developing the

API within your Flask app that you can

receive user inputs for model features,

load the model, make prediction, and

return the result. And after this, we'll

develop the UI. We'll design a

user-friendly interface. We'll create a

simple and intuitive UI that lets users

input the feature like location, size

and submit them for prediction. And

after that we'll move to the deployment

phase. We'll use a cloud platform like

Heroku, AWS or Google Cloud to deploy

your Flask application. And then we have

the maintenance and updates. We'll

monitor and update regularly for the

performance and use the model as needed

based on user feedback. So now moving to

the Python code and see how this model

can be created. So here we'll start

importing the libraries and module and

we are using flask ple and jsonify and

we will start with app initialization.

We'll initialize a new flask web

application that would be a special

variable which gives python files a

unique name to differentiate between

them when they are important into other

scripts. And after that we'll load the

model. So loading a pretend machine

learning model from the file system. So

this model is assumed to be saved in the

same directory as this script. So the

model is loaded in RB mode which stands

for read binary. And after that we'll

move to API route and prediction

function. So we will define an API

endpoint at predict that listens for

post request. This is the URL that the

front end of the web application will

call to send data to the back end. And

after that we'll start with predicting

the function. And here we have extract

features that retrieves data sent into

the JSON format from the post request

that is request get_json and the force

we have set it as true here and

forcefully formats the request data into

JSON ensuring compatibility and then

we'll extract the relevant features that

is location size and amenities from the

JSON object and store them in a list as

expected by the model and after

preparing the features we'll make the

prediction we'll use the loaded model to

make a prediction based bas on the

provided feature and then we have the

return prediction method. Here we will

convert the prediction result into JSON

format using JSON and send it back to

the client and this will ensure that the

response can be easily handled by the

client application. So this was all

about the question number 17. Now moving

to the question number 18 that is based

on analyzing. And now we move to the

question number 18 that is based on

analyzing geospatial data. And your

question is you are tasked with

analyzing geospatial data to help a city

improve its public transportation

system. The data includes GPS

coordinates of bus stops, ridership

numbers and traffic patterns. What steps

would you take to analyze this data? And

can you provide a sample Python code

snippet to illustrate how you might

visualize bus stop location and

ridership? So you can start answering

this question that juice better data

analysis can provide critical insights

into how effectively a public

transportation system serves its city

and guide improvements and there's an

detailed approach for this and we can

start with data preparation and in this

we'll do data collection and data

cleaning and after this step we'll move

to the next step that is explorative

data analysis and in this we'll have

statistical summary we'll generate

descriptive statistics and then we have

correlation analysis

And after moving that we have geospatial

visualization that is mapping bus stop.

We'll plot the locations of bus stop on

a map to visually assess their

distribution across the city. And after

that we have heat maps that will create

ridership data to identify hot sports

and areas with potential service gaps.

And after geospatial visualization we'll

move with spatial analysis. We have

proximity analysis that will analyze the

proximity of bus stop to key areas like

commercial centers or residential areas.

And now moving to the fifth step that is

optimization and recommendation. So

we'll have a route optimization that

will suggest modifications to route

based on traffic patterns and ridership

demand and the policy recommendations

that will provide actionable

recommendations for improving bus

frequencies. Now move to the sample

Python code where we can define this

model and use it accordingly. And here

we will start importing the libraries

and modules. And here we'll start with

importing geopandas and m lib dotpipo.

And after importing we'll start with

data loading. So we will declare a

variable bus stops and load the bus

stops data from a shape file. So shape

files are popular geospatial vector data

formats for geographic information

system software and then we have the

wrership that will load wrership data

from a CSV file which includes columns

for longitude latitude and ridership

levels and after that we'll create geo

data frame that will convert the

wrership data frame into a geo data

frame and this step involves creating a

geometry column from the longitude and

latitude columns and then we have the

plotting one here we will plot the

graphs that would figures and axis and

create a figure for the single subplot

with a specified size that is 10 + 10

in. And then we have city map.plot. It

is assumed that there is a base map of

the city loaded as a geo data frame

named city map. This is plotted first

with a light gray color to serve as a

background for the other layers. So this

was all about question number 18. Now

moving to question number 19 that is

based on predictive maintenance using

machine learning and the question is you

are tasked with developing a predictive

maintenance system for a manufacturing

plant that relies heavily on automated

machinery. So the data available

includes machine operational parameters,

maintenance history and failure

incidents. What steps would you take to

develop a predictive model and can you

provide a sample Python code? So you can

start with predictive maintenance that

is essential in manufacturing as it

helps prevent equipment failures

reducing downtime and maintenance cost.

And here you would have a detail

approach or predictive model for this

starting with data collection and

integration. Then you can do EDA that is

exploratory data analysis and then we

can perform feature engineering and then

move to data prep-processing task and

then the selection model and training

and after that we have model evaluation

and deployment technique that we can do

for the model using appropriate metrics

such as precision, recall and F1 score.

So this was all about question number

19. So now move to question number 20

that is based on personalization using

machine learning and your question is

you are tasked with developing a machine

learning model to personalize content

recommendations for users on a media

streaming platform. The data available

includes user demographic retails

viewing history and ratings. So what

steps would you take to build a model

for personalized recommendations and can

you provide a sample Python code for

that? So you can start answering this

with creating a personalized

recommendation systems. This would be

essential for engaging users by

providing content that is relevant to

their interest. And there will be a

systematic approach or personalized

content recommendation. We'll start with

data collection and integration. And

after that, we'll perform EDA that is

explorative data analysis. And then we

have feature engineering. In this we'll

interact features and the temporal

features. We'll include time based

features to capture trends and

seasonality in viewing behavior. And

then we'll select the model that is by

collaborative filtering and hybrid

models. And then we'll train the model

and validation and implement and monitor

them. And after that we'll deploy the

model. So let's start with beginner

level questions. And number one is what

is machine learning? So machine learning

is a subset of artificial intelligence

that involves the use of algorithms and

statistical models to enable computers

to perform task without explicit

instructions. that is by relying on

patterns and interference. And now

moving to number second question that is

what are the different types of machine

learning. So the three main types of

machine learning are number one is

supervised learning and then comes

unsupervised learning and then there is

reinforcement learning. Now moving to

next question that is third that is what

is supervised learning. So supervised

learning involves training a model on a

label data set which means each training

example is paired with an output label.

The model learns to predict the output

from the input data. Now moving to the

fourth question that is what is

unsupervised learning. So unsupervised

involve training a model on data that

does not have labeled responses. The

model tries to learn the patterns and

the structure from the input data. So

guys, these are the beginner level

questions and now we'll move to the

fifth question that is what is

reinforcement learning. So reinforcement

learning is a type of machine learning

where an agent learns to make decisions

by performing actions and receiving

rewards or penalties. The goal is to

maximize the cumulative reward. So now

moving to the sixth question that is

what is a model in machine learning. So

a model in machine learning is a

mathematical representation of a real

world process. It is trained on data to

recognize patterns and make predictions

or decisions based on new data. So now

moving to seventh question that is what

is overfitting? So overfitting occurs

when a machine learning model performs

well on the training data but poorly on

new unseen data. It indicates that the

model has learned the noise and details

in the training data instead of the

actual patterns. So now coming to

question number eight that is what is

underfitting? So underfitting occurs

when a machine learning model is too

simple to capture the underlying

patterns in the data. It performs poorly

on both the training data and new data.

Now move to the next question that is

ninth question and the question is what

is a confusion matrix? So confusion

matrix is a table used to evaluate the

performance of a classification model.

It summarizes the number of correct and

incorrect predictions made by the model

and that is categorized by each class.

Now moving to the 10th question that is

what is cross validation? So cross

validation is a technique for assessing

how the results of a statistical

analysis will generalize to an

independent data set. It involves

partitioning the data into subsets.

Training the model on some subsets and

validating it on the remaining subsets.

This was all about that is the 10th

question or the overall 1 to 10

questions for beginner level. Now we'll

move to intermediate level and here

we'll cover 10 questions. So we'll start

with 11th question that is what is a ROC

curve. So ROC that is receiver operating

characteristic curve. It is a graphical

representation of a classifier's

performance across different thresholds.

It plots the true positive rate that is

TPR against a false positive rate that

is FPR. Now moving to 12th question that

is what is precision and recall. So

precision is the ratio of correctly

predicted positive observations to the

total predicted positives and recall is

the ratio of correctly predicted

positive observations to all actual

positives. So the formula is precision

equal to TP/TP

plus FP and the recall is TP/TP

+ F_sub_1. So now we'll move to the 13th

question that is what is the F1 score?

So the F1 score is the harmonic mean of

precision and recall. It provides a

balance between the two metrics and is

useful when you need to balance

precision and recall. F1 score is equal

to twice into precision into recall and

that is divided by precision plus

recall. Now we'll move to 14th question

and here we will cover regularization.

So the question is what is

regularization? So it is a technique

used to prevent overfitting by adding a

penalty to the model's complexity and

the common types of regularization

include L1 that is lasso and L2 ridge

regularization. Now we'll move to the

15th question that is what is the bias

variance tradeoff. So the bias variance

trade-off is a fundamental issue in

machine learning that involves balancing

the error introduced by the model's

assumptions and the error due to model

complexity. So a good model should have

low bias and low variance. Now we'll

move to the question number 16 that is

what is feature engineering. So feature

engineering is the process of creating

new features or modifying existing ones

to improve the performance of a machine

learning model. It involves techniques

like normalization and coding

categorical variables and creating

interaction terms. So now we'll move to

question number 17 and that is about

gradient descent. So the question is

what is gradient descent and your answer

is gradient descent is an optimization

algorithm used to minimize the cost

function in machine learning models and

it iteratively adjust the model

parameters in the direction of the

steepest descent of the coast function.

So with this we'll move to the 18th

question and that will cover with the

difference between bagging and boosting.

So the question is what is difference

between bagging and boosting and you

could answer this with starting with

bagging that is bootstrap aggregating

that involves training multiple models

on different subsets of the data and

averaging their predictions. Then comes

boosting that involves training models

sequentially with each new model

focusing on correcting the errors of the

previous ones. And then we have the

question number 19 that is what is a

decision tree? So a decision tree is a

nonparametric supervised learning

algorithm used for classification and

regression. It splits the data into

subsets based on the value of input

features resulting in a treel like

structure of decisions. Now we'll move

to question number 20 that is what is a

random forest. So random forest is an

ansemble learning method that combines

multiple decision trees to improve the

accuracy and robustness of the model. It

builds each tree using a random subset

of features and data points and then

averages their predictions. So these

were the questions that are for the

intermediate level and these are just

the basic questions or I will just say

the theoretical questions that can be

asked in an interview. So be prepared

for that. Now we'll move to the advanced

level interview questions and we'll

start with question number 21. And here

also we'll cover the 10 questions. So

number one question or that is 21

question and the question is what is a

support vector machine? So support

vector machine is a supervised learning

algorithm used for classification and

regression. It finds the optimal hyper

plane that maximizes the margin between

different classes in the feature space.

And then comes question number 22 that

is what is principal component analysis.

So principal component analysis is a

dimensionality reduction technique that

transforms highdimensional data into a

lower dimensional space by finding the

directions that is principal components

that maximize the variance in the data.

And then comes the question number 23

that is what is a neural network? So a

neural network is a series of algorithms

that attempt to recognize underlying

relationships in a set of data through a

process that mimics the way the human

brain operates. It consists of layer of

interconnected nodes or neurons. And

then comes the question number 24 that

is what is deep learning? So deep

learning is a subset of machine learning

that involves neural networks with many

layers that is deep neural networks and

it is particularly effective for task

like image and speech recognition. Now I

move to question number 25 that is what

is convolutional neural network that is

CNN. So we will start the answer by

answering the interviewer that a

convolutional neural network is a type

of deep learning model specifically

designed for processing structured grid

data like images. It uses convolutional

layers to extract special features or

the spatial features and patterns from

the input data. Now we move to the

question number 26 that is what is a

recurrent neural network or RNN. So a

recurrent neural network is a type of

neural network designed for sequential

data and it has connections that form

directed cycles allowing it to maintain

a memory of previous inputs and process

sequences of data. So this was all about

question number 26 and now we will cover

the question number 27 that is what is

the difference between batch gradient

descent and stoastic gradient descent.

So batch gradient descent computes the

gradient of the coast function using the

entire training data set while

stochastic gradient descent that is SGD

computes the gradient using only one

training example at a time. So SGD is

faster but noisier. Now we move to

question number 28 that is what is

dropout in neural networks. So dropout

is a regularization technique used in

neural networks to prevent overfitting

and it involves randomly setting a

fraction of the neurons to zero during

training forcing the network to learn

more robust features. And now we'll move

to question number 29 and that will be

about transfer learning. And your

question is what is transfer learning?

So we'll answer this to the interviewer

by starting that transfer learning is a

technique in machine learning where a

model developed for one task is reused

as the starting point for a model on a

second related task. It is particularly

useful when there is limited data

available for the second task. Now we'll

move to the last question and the 30th

question. So that is what is a

generative adversial network that is GN.

So you can start answering this. So

generative adversial network is a type

of deep learning model consisting of two

neural networks a generator and a

discriminator that are trained

simultaneously. The generator creates

fake data while the discriminator tries

to distinguish between real and fake

data leading to the generator producing

increasingly realistic data. And these

questions and answers are over and these

covers a wide range of topics in machine

learning and should help prepare for

interviews at various levels.

>> And with that we have reached the end of

the session on the AI and machine

learning engineer full course for

beginners. If you have any doubts or

questions about this video let us know

in the comment section below and a team

of experts will be happy to help you.

Until next time thank you and keep

learning. Stay tuned for more from

SimplyLearn.

Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.