Afrikaans
Akan
Albanian
Amharic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranî)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
So in this lecture, we are going to discuss some common at times serious transformations, if you're
familiar with machine learning, then you know that it's often useful to transform your data before
passing it into a machine learning model.
For example, standardization or min scaling four time series, we'll be discussing three common transformations
the power transform, the log transform and the Buzzcocks transform.
As you'll see, these all essentially serve the same purpose.
So let's start with the power transform, the power transform involves raising all your data points
to a power, for example, by raising every data point to the power of one half, you'll be taking the
square root of your data set.
So why is this useful?
Well, imagine that your data appears to grow quadratic in time.
If you take the square root, the result would be that you transform your data to grow linearly.
So why is that useful?
Well, you'll soon learn about some machine learning models that can learn linear trends very well,
but there's no model for quadratic trends or Kubic trends and so forth.
Thus, by transforming your data to appear like it has a linear trend, you give your model a better
chance of forecasting future data points and modeling the true nature of the Time series more closely.
So another transformation with a similar purpose is the log transform, like the power transform, it
basically ends up squashing your data into a smaller range.
In fact, a lot of the time I'll just end up using the log transform by default without considering
other options.
One common application of the log transform is in finance and finance.
It's common to model stock prices as following a normal distribution.
It's also common to model log returns instead of returns based on percentages.
As an example, this is the basis for the famous Black-Scholes formula.
Note that one possible issue with the log transform is that it doesn't accept zero or negative values
as input.
For this reason, it can only be used for data which is strictly positive for data that might be non-negative.
It's common to simply add one before taking the log.
OK, so a third transform we're going to discuss is the box cox transform, which generalises the concept
of both the power transform and the log transform.
You can see that it involves this parameter lambda, which is the power to use when taking the transform.
So why does this make sense?
This makes sense because the natural logarithm is actually the limit of this specific power transform
as the power approaches zero.
Now inside the box Cox function will automatically choose the value of Lambda for us.
So we don't need to worry about finding the optimal value ourselves.
But if you're interested in learning how this value is chosen, I'd encourage you to check out the CPA
documentation as well as this article I've included in extra reading tea.
So one common reason people give for why they use the Buzzcocks transform is that they want to make
the data normally distributed.
However, note that this motivation does not apply to Raw Time series.
So why is this?
Well, remember that Time series data is dynamic.
It changes in time.
It can have a trend.
So when you take time series data and plot a histogram hoping that it will be normal, this is actually
the wrong thing to do was discuss this more later in the course.
But in order to take data over time and plot its distribution or histogram, we need that data to be
stationary.
Stationary essentially means distribution doesn't change over time.
So why is this a requirement?
Well, imagine you have some data which simply follows a line that grows at a constant rate.
Does plotting the histogram of this data makes sense?
The answer is no.
What do we want this to be normally distributed?
The answer is no.
In fact, this data behaves much better with a linear trend.
The point of plotting a histogram is to understand the distribution of the data, but the distribution
at the bottom of this plot is clearly different from the distribution at the top of this plot.
Therefore, it makes no sense to mix this data together into a single histogram.
This does not tell us how the data is distributed.
The final topic I want to discuss in this lecture is why the log transform is deeply fundamental.
Not only is it useful mathematically, but it also seems to be part of nature itself.
One example of this is perception.
For example, although a normal conversation is ten thousand times louder than a whisper, it doesn't
have ten thousand times the effect on your senses.
That's why we use the decibel scale to measure sound, which is essentially a log transform.
Another example of how the logarithm seems to simply be a part of nature is how we as humans interpret
numbers.
For example, if you have one thousand dollars in the bank, then losing one thousand dollars would
be a pretty big deal.
But if you have one billion dollars in the bank, spending one thousand dollars on a pair of jeans would
feel completely normal.
Another way to think of this is imagine going from zero dollars in wealth to one million.
That's a pretty big jump.
How about one million to two million?
Although you still made the same amount of money, its utility is less so.
One might model the utility of wealth as the logarithm of the wealth and.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.