Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranî)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
The objectives of this demo are to describe common data
challenges, describe the purpose of the data warehouse, and
to also describe the data warehouse ecosystem.
Today, like in many organizations,
there are numerous source systems.
These systems have been designed to capture operational
workloads, be them Sale systems, HR systems, finance systems
that could also be cloud based, Data extracts, massive amount of
extracts like Web logs that are sitting on big data stores.
Somehow users need to connect to and access this data.
And therein is a challenge for operational reports,
one approach is, that you could connect direct to these systems.
In fact, this is a supportive approach when you consider
a Sale system and the need to raise an invoice, this is
a report that is driven from the Operational data system.
But let's consider the other requirements of our users.
Analytics, the need to aggregate,
summarize, drill through this type of activity on data,
might seems simple from an interface perspective but
is quite demanding across the systems.
For example,
looking at a pivot table that it shows me employees on the rows,
the months of the year on columns and at the intersection,
I see the sum of sales sold by each employee by month.
That looks simple but what may not be clear to you
is that the underlying Source System is storing
billions of rows of data that needed to be retrieved,
filtered, grouped, aggregated simply to produce that result.
That could be intensive and especially for
a Source System that has been optimized for
write intensive activities, not read intensive.
Dashboards, Alerting systems,
Scorecards that compare goals to actuals.
These are all requirements that are not well driven through
Operational Reporting direct from Source Systems.
So, what could the solution be?
Well, one approach might be to reduce contention, is that you
might replicate these systems and that could be achieved, for
example with SQL server, through replication or
through database mirroring, both high availability strategy and
these mirrors could also be read only replicas,
allowing us to perform reporting.
However, while this might reduce contention, the structures and
typically of relational databases with operational
workloads to insert data are efficiently are highly
normalized, and this doesn't work so well for the analytic
requirements that we might want to drive from this.
So what we could consider is introducing this
data storage and aggregation.
A very generic box here.
And let's take a look and build up some scenarios
of how we could implement more effective User Access and
Business Intelligence driven from optimized data stores.
Here's an example of details in an Operational Data Store.
This is referred to as an ODS.
And where there's a need for Operational reporting and real
time up to date data without impacting on the Source Systems,
and what they've been designed to do.
An ODS could on a very frequent basis, collect and
integrate data from Operational systems and
support Operational reporting.
The focus however on this course is more
about the delivery of the Enterprise Data Warehouse.
And we're gonna talk in this course about different
architectures, and we see one here which is the Enterprise
Data Warehouse consisting all the series of Data Marts.
And now, I'll describe these as subject specific store like
Sales, Operations, HR and Finance.
And these still relational sources are optimized for
analytic work loads.
They're optimized for read intensive operations.
So, the question then is,
well, how do we get data from the Source Systems
loaded into these and a Enterprise Data Warehouse
structures to support User Access?
Well, the first discussion will be commonly
with large implementation is to introduce a staging area.
And the staging area in relational format is a place
to land Operational Data.
Often in our design approach, we wanna get in and out as quickly
as possible minimizing the impact on the Source Systems.
So with read only access perhaps the logic is something like
this, retrieve all Sales Data, since the last
time you retrieved it until now, i.e.the last 24 hours.
So truncate the staging, load in the incremental transactions
that have taken place since the last ETL process and
then the staging system supports interrogation,
transformation, cleansing it also supports restartability,
if there was a need to redo detail process.
Now, supporting the staging system could be a Master Data
System, where there's an identified need to maintain for
consistency purposes certain business entities, for example,
products, geography.
We might want to maintain golden records of data that are not
possible to maintain directly in our Operational system maybe
because it doesn't support it, there's no interface, or perhaps
products are actually defined in multiple systems, and so we have
no single place to take as an authoritative store of data.
So, Master Data System can solve many of those challenges.
The next consideration is in relation to garbage in,
garbage out.
If you collect garbage data or sourced garbage data from your
systems and load this into the Data Warehouse
you cannot expect quality decisions to be made, so
there may be data quality cleansing systems.
Knowledge basis on how to correct and
standardize or even repair or deduplicate data.
So collectively, Staging Systems,
Master Data Reference Systems and Data Quality Systems
can be used to drive a periodic ETL process, that according to
the business rules that define good quality data and
ETL process extract, transform, and load can periodically
load from this systems and load into the Data Warehouse.
Once loaded into the Data Warehouse it is clean,
consistent, credible current data that is available for
production reporting.
Now this data Marts are still relational data bases.
And we just mentioned that as Source System,
they're not the most efficient source to retrieve from.
This is because they often designed in third normal form
and that is a design optimization for
write intensive operations.
As relational databases, but with an analytic workload.
We still design in terms of tables, columns, and
relationships, but we use different and
mature methodologies that support the analytic workloads.
Dimensional modeling is the topic here, and
you may well be familiar with fact tables, dimension tables.
Now, relational systems, even when they're designed optimally
in this fashion, they're still inherently slow.
And so, what you will find in a Data Warehouse are data models
like the Sales Model and the Finance Model here.
These can be referred to as cubes, or
BI semantic models, data models whatever you name them, they're
essentially a very convenient access point for your end users.
Your end users, granted permission,
can connected to these models.
They can work with high performance queries and
analytic query workloads.
This is achieved often because these data models may cache and
place in memory or vision structure on disk in the memory,
and what is enable is very high performance slicing and
dicing very natural to answer the type of analytic
questions that business typically has.
Now, these data models are also a great place to encapsulate
business logic even difficult calculations, time manipulation
can be encapsulated far easier here and can be achieved with
the logic available in relational querying.
In addition, there are other great things,
like translations for different languages, actions to support
moving from the data to other experiences like reports or
drill through data sets.
And lastly, there's the concept of security.
When you have different permission sets for
different audiences,
it's quite difficult in a relational system to apply this.
And yet, data models have roles and ways to define permission
sets that can be quite complex right to a granular level that
different people can see different data.
Now, also at this level,
you see on the presentation the Churn Analysis Model.
This is, in fact, a machine learning or even a data mining
model that has been processed against the Data Marts
looking for patterns of interest, in this case, it might
be looking for characteristics of customers that leave you.
And if you can identify these, that's useful.
So they provide exploration, and beyond that,
we do trust the patterns that they have surfaced.
They can be deployed as predictive models.
And they can be used in reporting and analytics or
to drive other business functionality.
The last build then of this slide is to introduce from
a User Access perspective, the need for Self Service BI.
We should never recognize that an Enterprise Data Warehouse
will deliver 100% of the business requirements.
I might strive somewhere between 80 and 90%.
Yet as the business evolves, you'll find that
the Data Warehouse is a major undertaking, and
it is not so agile to simply adapt
quickly to the new questions that the business might raise.
So, Self Service Business Intelligence is a way to fill
that gap by empowering the right people in the organization with
tools and access to data and training.
They can connect to the Data Warehouse resources,
even the data marts directly or even the models.
And they may construct new models by extending or
adding new logic beyond what the Data Warehouse delivers.
This is a valid form of BI, and we should see it as a mutual
benefit to the organization by extending and working with
the resources already deployed in an Enterprise Data Warehouse.
This then end to end describes the business case for
the Data Warehouse in the organization in today.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.