Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranî)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
This is how intelligence is made.
A new kind of factory,
generator of tokens, the building blocks
of AI.
tokens have opened a new frontier,
turning data into knowledge and drawing
on all we have learned.
Tokens are harnessing a new wave of
clean energy
and unlocking the secrets of the stars.
In virtual worlds, they help robots
learn and in the physical world perfect,
forging new paths
and clearing the way for a bountiful
harvest.
In the moments that matter, tokens are
already there.
And in the miles between, they never
stop.
They work where human hands cannot.
So we may all breathe easier.
And the smallest hearts beat stronger.
Tokens are helping us break new ground
on a scale never attempted
to empower the world.
So we can reach star cloud one.
Separation confirmed. well beyond it.
Together we take the next great leap
into a bright new future
built for all mankind.
And here
is where it all begins.
Welcome to the stage, Nvidia founder and
CEO, Jensen Wong.
Welcome to GTC.
I just want to remind you this is a tech
conference.
All these people lining up so early in
the morning. All of you in here, it's
great to see you.
GTC
GTC. We're going to talk about
technology. We're going to talk about
platforms. Nvidia has three platforms.
You think that we mostly talk about one
of them. It's related to CUDA X. Our
systems is another platform and now we
have a new platform called AI factories.
We're going to talk about all of them
and most importantly we're going to talk
about ecosystems. But before I start,
let me thank our pregame show hosts. I
thought they did a great job. Sarah Go
of Conviction,
Alfred Lyn, Sequa Capital, Nvidia's
first venture capitalist, Gavin Baker,
Nvidia's first major institutional
investor. These three people are deep in
technology, deep in what's going on and
of course they have just a really broad
reach of technology ecosystem. And then
of course all of the VIPs that I hand
selected to join us today, allstar team.
I want to thank all of you for that.
I also want to thank all the companies
that are here.
Nvidia as you know is a platform
company. We have technology, we have our
platforms, we have re rich ecosystem and
today there are probably 100% of the
hundred trillion dollars of industry
here. 450 companies sponsored this
event. I want to thank you. A thousand
technical sessions, 2,000 speakers. This
is this conference is going to cover
every single layer of the five layer
cake of artificial intelligence from
land power and shell the infrastructure
to chips to the platforms the models and
of course the most important and
ultimately what's going to take get this
industry taken off is all of the
applications.
What it all began it all began here.
This is the 20th anniversary of CUDA.
We've been working on CUDA for 20 years.
For 20 years, we've been dedicated to
this architecture. This revolutionary
invention, SIMT, single instruction,
multi-threaded, writing scalar code
could spawn off into multi-threaded
application. much much easier to program
than CINDI. We recently added tiles so
that we could help people program tensor
cores and the structures of mathematics
that are so foundational to artificial
intelligence today.
Thousands of tools and compilers and
frameworks and libraries
in open source. There's a couple of
hundred thousand public projects. CUDA
literally is integrated into every
single ecosystem.
This chart
basically describes 100% of Nvidia's
strategies. You've been watching me talk
about this slide from the very
beginning. And ultimately, the single
hardest thing to achieve is the thing on
the bottom, installed base. It has taken
us 20 years to now have built up
hundreds of millions of GPUs and
computing systems around the world that
run CUDA. We are in every cloud. We're
in every computer company.
We serve just about every single
industry. The installed base of CUDA is
the reason why the flywheel is
accelerating. The install base is what
attracts developers who then creates new
algorithms that achieves a breakthrough.
For example, deep learning. There are so
many others. Those breakthroughs leads
to entirely new markets which builds new
ecosystems around them with other
companies that join which creates a
larger installed base. This flywheel,
this flywheel is now accelerating. The
number of downloads of Nvidia libraries
is incredibly accelerating. is at a very
large scale and growing faster than
ever. This flywheel is what makes this
computing platform able to sustain so
much applications, so many new
breakthroughs. But most importantly,
it also enables
these infrastructures to have
extraordinarily useful life. And the
reason for that is very obvious. There's
so many applications that you can run on
Nvidia CUDA. We support the entire every
single phase of the AI life cycle. We
address every single data processing
platform. We accelerate scientific
principled solvers of all different
kinds. And so the application reach is
so great that once you install Nvidia
GPUs, the useful life of it is
incredibly high. It is also one of the
reasons why Ampear that we shipped them
some six years ago the pricing of Ampear
in the cloud is going up. And so all of
that is made possible fundamentally
because the install base is high, the
flywheel is high, the developer reach is
great. And when all of that happens and
we continuously update our software,
the computing cost
declines. The combination of accelerated
computing speeding up applications
tremendously. Meanwhile, as we continue
to nurture and continue to update
software over its life, not only do you
get the first time pop, you get the
continuous cost reduction of accelerated
computing over time. And we're willing
to nurture, willing to support every
single one of these GPUs in the world
because they're all architecturally
compatible. We're willing to do so
because the install base is so large. If
we release a new optimization, it
benefits millions.
This applies to everybody in the world.
This combination of dynamics is what
makes the NVIDIA architecture expand its
reach, accelerating its growth, at the
same time driving down computing cost,
which ultimately
encourages new growth. So, CUDA is at
the center of it. But our journey that
could actually started 25 years ago.
GeForce.
I know how many of you grew up with
GeForce.
GeForce is Nvidia's greatest marketing
campaign.
We attract future customers starting
long before you could afford to pay for
it yourself. Your parents paid
Your parents paid your par your parents
paid for you to be Nvidia customers. And
every single year they paid up year
after year after year until someday you
became an amazing computer scientist and
became a proper customer, a proper
developer. But this is this is the house
that GeForce made 25 years ago. We
started our journey which led to CUDA.
25 years ago, we invented the
programmable shader. A perfectly
unobvious invention to make an
accelerator programmable. The world's
first programmable accelerator, the
pixel shader 25 years ago. That led us
to explore further and further 20 years
later, 5 years later, the invention of
CUDA. One of the biggest investments
that we made and we couldn't afford it
at the time. and it consumed the vast
majority of our company's profits was to
take CUDA on the backs of GeForce to
every single computer. We dedicated
ourselves to creating this platform
because we felt so much we felt so
strongly about its potential. But
ultimately the company's dedication to
it despite the hardships in the
beginning believing it every single day
for third for 13 generations or 20 years
we now have CUDA installed everywhere.
The pixel shader
led to of course the revolution of
GeForce.
And then 10 years ago, we introduced
about 10 years ago, what is it, eight
years ago, we introduced RTX,
a complete redesign of our architecture
for the modern era of computer graphics.
GeForce brought CUDA to the world.
GeForce
therefore enabled Alex Kruefky and Ilas
Susver and Jeff Hinton, Andrew Ang and
so many others to discover that the GPU
could be their friend in accelerating
deep learning. It started the big bang
of AI. 10 years ago, we decided that we
would fuse
programmable shading and introduce two
new ideas. ray tracing, hardware ray
tracing, which is incredibly hard to do.
And a new idea at the time, imagine
about 10 years ago, we thought that AI
would revolutionize computer graphics.
Just as GeForce brought AI to the world,
AI is now going to go back and
revolutionize how computer graphics is
done all together. Well, today I'm going
to show you something of the future.
This is our next generation of graphics
technology. We call it neuro rendering.
The fusion,
the fusion of 3D graphics and artificial
intelligence. This is DLSS 5. Take a
look at it.
Heat. Heat.
Heat. Heat.
Is that incredible?
Computer graphics comes to life. Now
what did we do? We fused
controllable 3D graphics. The ground
truth of virtual worlds, the structured
data, remember this word, the structured
data of virtual worlds, of gener
generated worlds. We combine 3D
graphics, structured data with
generative AI,
probabilistic computing. One of them is
completely predictive, the other one
probabilistic yet highly realistic. We
combine these two ideas. Combine these
two ideas controlled through structured
data controlled perfectly and yet
generating at the same time. And as a
result,
the content is beautiful, amazing, as
well as controllable. This concept of
fusing structured information and
generative AI will repeat itself in one
industry after another industry after
another industry. Structured data is the
foundation of trustworthy AI. Well,
this is going to scare you a little bit.
I'm going to flip the slide. and don't
gasp.
So, we're going to go through the
schematic for the rest of the time.
This is my best slide. Every time I I
asked my I asked the team, "What's my
best slide?" Repeatedly, this was it.
They say, "Don't do it, Jensen. Don't do
it." I said, 'N no, this these seats are
free
for some of you.
So this is your price of admission. So
this is this is structured data. You've
heard of it. SQL, Spark, Pandas, Velox,
some of these really really important
very large platforms. Snow, snow, uh,
snowflake, data bricks, EMR, Amazon EMR,
um, Azure, Fabric,
Google Cloud, BigQuery. All of these
platforms are processing data frames.
These data frames are giant spreadsheets
and they hold all of life's information.
This is the structured data, the ground
truth of business. This is the ground
truth of enterprise computing. Well, now
we're going to have AI use structured
data and we better accelerate the living
daylights out of it. It used to be okay
and we would, you know, of course we
would accelerate uh structured data so
that we could do more. We could do it
more cheaply. We could do it more
frequently per day and keep the company
running at a much more synchronized way.
However, in the future, what's going to
happen is these data structures are
going to be used by AI and AI is going
to be much much faster than us. Future
agents are going to use structured
databases as well. And then of course
the unstructured database, the
generative database. This database is
represents the vast majority of the
world. Vector databases, unstructured
data, PDFs, videos, speeches, all of the
world's information. About 90% of what's
generated every single year is
unstructured data. Until now, this data
has been completely useless to the
world. We read it, we put it into our
file system, and that's it.
Unfortunately, we can't query it. We
can't search for it. It's hard to do
that. And the reason for that is because
there's no easy indexing of unstructured
data. You have to understand its
meaning, its purpose. And so now we have
AI do that just as AI was able to solve
multi-modality
perception you can and understanding you
can use that same technology
multimodality perception and
understanding to go read a PDF to
understand its meaning and from that
meaning embedded into a larger structure
that we can search into we can query
into. NVIDIA created two foundational
libraries. Just like we created RTX for
3D graphics, we created QDF for data
frames, structured data. We created QVS
for vector stores, semantic data,
unstructured data, AI data. These two
platforms are going to be two of the
most important platforms in the future.
super excited to see its adoption
throughout the network, this complicated
network of the world's data processing
systems. And the reason for that is
because data processing has been around
a long time and therefore so many
different companies and platforms and
services. It has taken us a long time to
integrate deeply into this ecosystem.
I'm super proud of the work that we're
doing here. And then today we're
announcing several of them. IBM
the inventor of SQL
one of the most important domain
specific languages of all kind of all
time is accelerating Watson X data with
KUDF let's take a look at it
60 years ago IBM introduced the system
360
the first modern platform for
generalpurpose computing launching the
computing era then SQL a declarative
language to query data without requiring
the computer to be instructed step by
step
and the data warehouse. Each the
foundations of modern enterprise
computing. Today, IBM and NVIDIA are
reinventing data processing for the era
of AI by accelerating IBM Watson X. Data
SQL engines with NVIDIA GPU computing
libraries. Data is the ground truth that
gives AI context and meaning. AI needs
rapid access to massive data sets.
Today's CPU data processing systems
can't keep up. Nestle makes thousands of
supply chain decisions every day. Their
order to cache data mart aggregates
every supply order and delivery event
across global operations in 185
countries.
On CPUs, Nestle refreshed the data mart
a few times a day. With accelerated
Watson X data running on Nvidia GPUs,
Nestle can run the same workload five
times faster at 83% lower cost.
The next computing platform has arrived.
Accelerated computing for the era of AI.
NVIDIA accelerates data processing in
the cloud. We also accelerate data
processing on prem. As you know, Dell is
the worldleading computer systems maker
and they also are one of the world's
leading storage providers and they
worked with us to create the Dell AI
data platform that integrates QDF and
QVS to create an accelerated data
platform. well for the era of AI and uh
this is an example of what they did with
NT data huge speed up this is cloud
Google cloud and Google cloud as you
know we've been working with Google
cloud for a very long time we accelerate
Google's vertex AI we now accelerate
bigquery really important uh framework
and really important platform and this
is an example of our work together with
Snapchat where we reduce their cost of
computing by nearly 80%.
When you accelerate data processing,
when you accelerate computing, you get
the benefit of speed, you get the
benefit of scale, but most importantly,
you also get the benefit of cost. And so
all of those come together as one. It
was originally called Moore's law.
Moore's law was about getting
performance doubling every couple of
years. It's another way of saying so
long as the price remains about the same
and most computers remained about the
same, you're also getting twice the
performance every year or you're
reducing the cost of computing every
single year. Well, Moore's law has run
out of steam. We need a new approach.
Accelerated computing allows us to take
these giant leaps forward and as you
will see later because we continue to
optimize the algorithms
and Nvidia is an algorithm company. As
we continue to optimize the algorithms
and because our our reach is so large
and our install base is so large we can
reduce the computing cost increasing the
scale increasing the speed for everybody
continuously. This is Google cloud. You
could see this pattern I just mentioned.
I just wanted to show you three versions
of it. Nvidia built the accelerated
computing platform has a bunch of
libraries on top. I gave you three
examples. RTX is one of them. QDF is
another. KVS and we'll show you a few
more. These libraries sit on top of our
platform. But ultimately
we integrate into the world's cloud
services into the world's OEMs and
together and other platforms that I'll
show you together were able to reach the
world. This pattern Nvidia, Google
Cloud, Snapchat will repeat over and
over again. And kind of looks like this.
And so this is one example. Nvidia with
Google Cloud. We accelerate Vertex AI.
We accelerate Bitquery. We accelerate.
We're I'm super proud of the work that
we've done with Jackson XLA. We are
incredible on PyTorch. We're the only
accelerator in the world that's
incredible on PyTorch and incredible on
Jackson XLA. And the customers that we
support, the base 10s, the Crowd
Strikes, Puma, Salesforce, they're not
our customers, but they're customers,
developers of ours that we've integrated
the NVIDIA technologies into that we can
then land on the clouds.
Our relationship with cloud service
providers are essentially us bringing
customers to them. We integrate our
libraries, we accelerate workloads, and
we land those customers in the clouds.
And so, as you could see, most of our
cloud service providers love working
with us. And um they're always asking us
to land the next customer on their
cloud. And I just want to let you know
there are a lot of customers.
We're going to accelerate everybody. And
so, there will be lots and lots of
customers will be able to land in your
cloud. Just be patient with us. And so
this is Google Cloud. This is AWS. We've
been working with AWS a long time. And
one of the areas, one of the one of the
things I'm super excited about this year
is we're going to bring open AI to AWS.
And so it's going to drive enormous
consumption of cloud computing at AWS.
It's going to expand the reach, expand
the compute of open AI. And as you know,
they are completely compute constrained.
And so AWS, we accelerate EMR, we
accelerate SageMaker, we accelerate
Bedrock. NVIDIA's integrated really
deeply into AWS. They were our first
cloud partner,
Microsoft Azure.
NVIDIA's A100 supercomputer
um was the the first one we built was
for Nvidia. The first one we installed
was at Azure. And that led to the inter
the uh the big successful partnership
with open AI but we've been working with
Azure for quite a long time. We
accelerate Azure cloud now it's uh their
AI foundry we partner deeply with we
accelerate Bing search we work with them
on Azure regions. This is one of the
areas that is incredibly important as we
continue to expand AI throughout the
world. One of the capabilities that we
offer is confidential computing.
That in confidential computing, you want
to make sure that even the operator
cannot see your data. Even the operator
cannot touch or see your models.
confidential computing. Nvidia's GPUs is
the first ones in the world to do that.
It's now able to support confidential
computing and protected deployment of
these very valuable open AI models and
and anthropic models throughout clouds
and different regions and all because of
our conf confidential computing.
Confidential computing is super
important. And here's an example where
we have different customers that we work
with. Synopsis, a great partner of ours.
were accelerating all of their EDA and
CA workflows. And then we landed at
Microsoft Azure.
We were Oracle's first AI customer.
Most people would have thought we were
their first supplier. We were their
first supplier also, but we were their
first AI customer. I'm quite proud of
the fact that I explained AI clouds to
Oracle for the first time and we were
their first customer. Since then,
they've really taken off. We've landed a
whole bunch of our partners there. Core
Coher and Fireworks and of course very
famously open AAI
a great partnership with Core
Core. They're the world's first AI
native cloud. A company that was built
with only one singular purpose to
provision to host GPUs as the era of
accelerated computing showed up and to
host for AI clouds. They've got some
fantastic customers and they're growing
incredibly. One of the platforms that
I'm quite excited about is Palunteer and
Dell. The three of our companies have
made it possible to stand up a brand new
type of AI platform, the Palunteer
ontology platform and AI platform. And
we could stand up these platforms in any
country in any airgapped region
completely on prem, completely on site,
completely in the field. AI could be
deployed literally everywhere without
our confidential computing capability
without our ability to build the
endtoend system as well as offer the
entire
accelerated computing and AI stack from
data processing whether it's vectors or
structures all the way to AI it wouldn't
have been possible I wanted to show you
these examples
this is our special working relationship
with the world's cloud service providers
and many well all of them are here and I
get the benefit of seeing them during
boot tour and it's just so incredibly
exciting. I just want to thank all of
you for the hard work. What NVIDIA has
done is this and you're going to see
this theme over and over again.
Nvidia is vertically integrated the
world's first vertically integrated
but horizontally open company
and the reason that's necessary is very
simple. Accelerated
computing is not a chip problem.
Accelerated computing is not a systems
problem. Accelerated computing has a
missing word. We just never say it
anymore. Application acceleration.
You if I could make a computer run
everything faster, that's called a CPU.
But that's run out of steam. The only
way for us to accelerate applications
going forward and continue to bring
tremendous speed up, tremendous cost
reduction is through application or
domain specific acceleration. I dropped
that phrase in the in the front and
therefore it just became applica
accelerated computing and that is the
reason why Nvidia has to be library
after library, domain after domain,
vertical after vertical.
We are a vertically integrated computing
company. There is no other way. We have
to understand the applications. We have
to understand the domain. We have to
understand fundamentally the algorithms.
And we have to figure out how to deploy
the algorithm
in whatever scenario it wants to be
deployed. Whether it's a data center,
cloud, onrem, at the edge, or in a
robotic system. All of those computing
systems are different. And finally, the
systems and chips. We are vertically
integrated. What makes it incredibly
powerful and the reason why you saw all
the slides is because Nvidia is
horizontally open. We work and integrate
Nvidia's technology into whatever
platform you would like us to integrate
into. We offer you the software. We
offer you libraries. We integrate with
your technology so that we can bring
accelerated computing to everybody in
the world.
Well,
this GTC is really a great demonstration
of that. You know, most of the time,
most of the time you'll see me talk
about these verticals and I'll use some
examples, but in every single case,
whether it's automotive f by the way,
financial services, the largest
percentage of attendees at this GTC is
from the financial services industry.
I know. I I'm hoping it's developers,
not traders.
Guys,
here's here's
here's one thing I wanted to say. And so
in the audience represents Nvidia's
ecosystem upstream of our supply chain
and downstream of our supply chain. And
we work we think about our supply chain
upstream and downstream. And it's just
so exciting that
our entire upstream supply chain this
last year
irrespective of whether you're a 50 year
old company, we have 70 year old
companies. We have a 150 year old
company who are now part of Nvidia
supply chain and partnering with us
either upstream or downstream. And last
year
you had your record year, did you not?
Congratulations.
We're on to something here. This is the
beginning of something very, very big.
And so if you look at accelerated
computing, we've now set the computing
platform. But in order for us to
activate those computing platforms, we
need to have domain specific libraries
that solve very important problems in
each one of the verticals that we
address. You see us addressing every
single one of this. Autonomous vehicles,
our reach, our breadth, our impact.
Incredible. We have a track on that.
financial services. I just mentioned
algorithmic trading is going from
classical machine learning with human
feature engineering called quant the
quants did that to now supercomputers
studying massive amounts of data
discovering insight and discovering
patterns by itself and so this is going
through its deep learning and its
transformer moment healthcare is going
is going through their chap GPT moment
some really exciting work that we're
there we We have a great keynote track
here. We have a great keynote track.
Kimberly Pal is doing a great keynote
track um for healthcare. We're talking
about AI physics or AI biology for drug
discovery, AI agents for customer
service and support of diagnos diagnosis
and of course physical AI, robotic
systems. All these different vectors of
AI have different platforms that NVIDIA
provides. industrial we are completely
resetting and starting the largest
buildout of human history and most of
the world's industries building AI
factories building chip plants building
computer plants are represented here
today media and entertainment gaming of
course real time AI platform so that we
could translation and broadcast support
and live live games and live video
enormous amount of it will be augmented
with AI. We have a we have a platform
called hollow scan quantum there are 35
different companies here building with
us the next generation of quantum GPU
hybrid systems uh retail and CPG using
Nvidia for supply chain using creating a
gentic shopping systems
AI agents for customer support a lot of
work being done here $35 trillion
industry robotics $50 trillion industry
in manufacturing Nvidia has been working
in this area for a decade now building
three computers, the fundamental
computers necessary to build robotic
systems. We are integrated with working
with literally every single company that
we know of building robots. We have 110
robots here at the show. And then
telecommunications
about as large as the world's IT
industry about$2 trillion dollars. We
see of course base stations everywhere.
It's one of the world's infrastructures.
It was the infrastructure of the last
generation of computing. That
infrastructure is going to get
completely reinvented. And the reason
for that is very simple. That base
station which is
it does one thing which is base station
is going to be an AI infrastructure
platform in the future. AI will run at
the edge. And so lots of lots of great
um uh great uh discussion there. And our
platform there is called Aerial or AI
RAM. Big partnership with Nokia, big
partnership with T-Mobile and many
others.
At the core of our business,
everything that I just mentioned,
computing platforms, but very
importantly, our CUDA X libraries, our
CUDA X libraries is the algorithm, the
algorithms that Nvidia invents. We are
an algorithm company. That's what makes
us special. That what that's what makes
it possible for me to be able to go into
every single one of these industries,
imagine the future and have the world's
best computer scientists describe and
solve problems, refactor it, reexpress
it,
and turn it into a library. We have so
many I think we have at this show, we're
announcing a hundred 100 libraries,
something 70 libraries, maybe 40 models
and that's just at the show. We're
updating these all the time. We're
updating them all the time. The
libraries is the crown jewels of our
company. It is what makes it possible
for that platform, the computing
platform to be activated in service of
solving a problem, making impact. One of
the biggest, one of the most important
libraries that we ever created, coupn
CUDA deep neural networks. It completely
revolutionized artificial intelligence,
caused a big bang of modern AI. Let me
show you a short video about CUDA X.
20 years ago, we built CUDA, a single
architecture for accelerated computing.
Today, we've reinvented computing. A
thousand CUDA X libraries help
developers make breakthroughs in every
field of science and engineering.
CU opt for decision optimization.
CU litho for computational lithography.
CDSS for direct sparse solvers.
Coup equivariance for geometryaware
neural networks.
Aerial for AI ran.
Warp for differentiable physics.
pair of bricks for genomics.
At their foundation are algorithms and
they are beautiful.
Wow.
Heat. Heat.
Heat.
Heat.
Heat. Heat.
Heat.
Heat.
Everything you saw was a simulation.
Some of it was principled solvers,
fundamental physics solvers. Some of it
was AI surrogates, AI physical models
and some of it was physical AI robotics
models. Everything was simulated.
Nothing was animated. Nothing was
articulated. Everything was completely
simulated. That is what fundamentally
Nvidia does. It is through the
connection of understanding of the
algorithms with our computing platforms
that we're able to open up to unlock
these opportunities. Nvidia is a
vertically integrated computing company
with open
horizontal integration with the world.
So that's CUDA X. Well, just now you saw
a whole bunch of companies. You saw
Walmart and you know there's L'Oreal and
incredible companies established
companies JP Morgan and Ro and these are
companies in companies that have defined
society to today. Toyota is here. These
are some of the largest companies in the
world.
It is also true
that there's a whole bunch of companies
you've never heard of. These are
companies we call them AI natives. a
whole bunch of small companies. This the
list is gigantic. I can't I couldn't
this is just a little tiny tiny bit of
it. And um I I couldn't decide whether
to show you more or show you less. And
so I I made it so that you couldn't see
any
and and nobody's feelings are hurt.
However, inside this list are a bunch of
brand new companies. There are companies
like for example you might have heard a
couple of them open AAI anthropic but
there's a whole bunch of others there's
a whole bunch of others and they serve
different verticals
something happened in the last two years
particularly this last year we've been
working with the AI natives for a long
time and this last year it just
skyrocketed and I'll explain to you why
it happened these this industry has
skyrocketed $150 billion dollars of
investment into venture investment into
startups, the largest in human history.
This is also the first time that the
scale of the investments went from
millions of dollars, tens of millions of
dollars to hundreds of millions of
dollars and billions of dollars. And the
reason for that is this is the first
time in history that every single one of
these companies
needs compute and lots and lots of it.
They need tokens, lots and lots of it.
they're either need they're either going
to create and build and create tokens
and generate tokens or they're going to
integrate
add value tokens
that are available created by anthropic
and open AAI and others and so this
industry is different in so many
different ways but the one thing that is
very clear the impact that they're
making this the incredible value that
they're delivering already is quite
tangible AI natives
All because we reinvented computing.
Just like during the PC revolution, a
whole bunch of new companies were
created. Just as just as uh during the
internet revolution, a whole bunch of
companies were created and mobile cloud
a whole bunch of companies were created.
Each one of them had their own standards
and all we're talk about one of the
major standards is that just happened.
Incredibly important. And this
generation, we also have our own large
number of very, very special companies.
We reinvented computing. It stands to
reason there's going to be a whole new
crop of really important companies,
consequential companies for the future
of the world. The the Googles, the
Amazons, the Metas, consequential
companies that have come as a result of
the last computing platform shift. We
are now at the beginning of a new
platform shift. But what happened in the
last couple years? Well, we've been
watching, as you know, we've been
working on deep learning and working on
AI, the big bang of modern AI. We were
right there at the spot and we've been
advancing this field for quite some
time. But why the last two years? What
happened in the last two years? Well,
three things. Chat GPT of course started
the generative AI era. It's able to not
just understand, perceive, and
understand. It's able to also translate
and generate generation of unique
content. I showed you the fusion of
generative AI with computer graphics and
it brought computer graphics to life.
You guys just everybody in the world
should be using chat GPT. I know I use
it every single morning. Used it plenty
this morning. And so chat GPT was the
generative AI the era. The second by the
way generative generative computing
versus the way we used to do computing.
It's not it's generative AI is a
capability of software but it has
profoundly changed how computing is
done.
Computing used to be retrieval based now
it's generative. Keep that thought in
mind when I talk about certain things
and you'll realize why it is that
everything that we do is going to change
how computers are architected, how
computers are provided, how computers
are going to be built out and what is
the meaning of computing altogether
generative AI 2023 end of 22 2023 the
next reasoning AI 01
which and then took off with 03
reasoning allowed it to reflect, allows
it to think to itself, allowed it to
plan, break down, break down problems
and decompose a problem it couldn't
understand into steps or parts that it
could understand. It could ground itself
on research. 01 made generative AI
trustworthy and grounded on truth. That
caused Chad GPT to simply took off. And
that was a very, very big moment. the
amount of input tokens that was
necessary in order to produce and the
amount of output tokens it need it
generated in order to reason the model
was a little bit larger it you know of
course you could have much larger models
the model 01 was a little bit larger not
much larger but its input token usage
for context
and its output token for thinking
increased the amount of computation
tremendously then came quad code the
first agentic model. It was able to read
files, code, compile it, test it,
evaluate it, go back and iterate on it.
Cloud code has revolutionized software
engineering. As all of you know, 100% of
NVIDIA is using a combination of CL or
oftentimes all three of them. Cloud
code, codeex, and cursor all over
Nvidia. There's not one software
engineer today who is not assisted by
one or many AI agents helping them code.
Cloud code completely revolutionizes the
the new inflection and the for the first
time.
You don't ask a AI what,
where, when, how.
You ask it
create, do, build.
You ask it to use tools,
take your context, read files. It's able
to agentically break down a problem,
reason about it, reflect on it. It's
able to solve problems, and actually
perform tasks. An AI that was able to
perceive became an AI that could
generate. An AI that could generate
became an AI that could reason. An AI
that could reason now became an AI that
can actually do work. Very productive
work. The amount of computation in the
last two years, we know that everybody
in this room knows the computing demand
for NVIDIA GPU is off the charts. Spot
pricing is skyrocketing. You couldn't
find a GPU if you tried. And yet, in the
meantime, we're shipping GPUs out,
incredible amounts of it, and demand
just keeps on going up. There's a reason
for that. This fundamental inflection.
Finally, AI is able to do productive
work and therefore the inflection point
of inference has arrived.
AI now has to think. In order to think,
it has to inference. AI now has to do.
In order to do, it has to inference. AI
has to read. In order to do so, it has
to inference. It has to reason. It has
to inference. every part of AI
every time it has to think it has to
reason it has to do it has to generate
tokens it has to inference it's way past
training now it's in the in the field of
inference so the in the inference
inflection has arrived
at the time when the amount of tokens
the amount of compute necessary
increased by roughly 10,000 times now
when I combine these to the fact that
since in the last two years the
computing demand computing demand of the
work has gone up by 10,000 times and the
amount of usage
the amount of usage has probably gone up
by a hundred times.
People have heard me say I believe that
computing demand has increased by 1
million times in the last two years. It
is the feeling that we all have. It is
the feeling every startup has. It's the
feeling that OpenAI has. It's the
feeling that Anthropic has. If they
could just get more capacity, they could
generate more tokens. Their revenues
would go up. More people could use it.
The more advanced, the smarter the AI
could become. We are now at that
positive flywheel system. We have we
have reached that moment. The
inflection, the inference inflection has
arrived. Last year at this time, I said
that
where I stood at that moment in time, we
saw about
$500
billion dollars.
We saw$500 billion dollars
of very high confidence demand and
purchase orders
for Blackwell and Reuben through 2026.
I said that last year.
Now, I don't know if you guys feel the
same way, but $500 billion is an
enormous amount of revenue.
Not one impressed.
I know why you're not impressed. Because
all of you had record years.
Well, I'm here to tell you
that right now where I stand, a few
short months after GTCDC,
one year after last GTC, right here
where I stand,
I see through 2027
at least $1 trillion
Now, does it make any sense?
And that's what I'm going to spend the
rest of the time talking about. In fact,
we are going to be short. I am certain
computing demand will be much higher
than that. And there's a reason for
that. So, the first thing is
um we did a lot of work in the last
year. Of course, as you know, 2025 was
NVIDIA's year of inference. We wanted to
make sure that not only were we good at
training and post- training, that we
were incredibly good at every single
phase of AI so that the investments that
were made, investments made in our
infrastructure could scale out for as
long as they would like to use it. And
the useful life of Nvidia's
infrastructure would be long and
therefore the cost would be incredibly
low. The longer you could use it, the
lower the cost. There's no question in
my mind Nvidia systems are the lowest
cost infrastructure you could get for AI
infrastructure in the world. And so the
first part was last year was all about
AI for inference and it drove this
inflection point. Simultaneously
we were very pleased last year that
Anthropic has come to Nvidia that MSL
Meta SL has chosen Nvidia and meanwhile
meanwhile and as a collection as a group
this represents
onethird of the world's AI compute open-
source models open-source models have
reached near the frontier and it is
literally everywhere and Nvidia as you
know today we're the only platform in
the world today that runs every single
domain of AI
across every single one of these AI
models
in language and biology and computer
graphics computer vision and speech
proteins and chemicals robotics and
otherwise edge or cloud any language
NVIDIA's architecture is funible for all
of that and we're incredible for all of
that. That allows us to be the lowest
cost, the highest confidence platform
because when you're building these
systems, as I mentioned, a trillion
dollars is an enormous amount of
infrastructure. You have to have
complete confidence that the trillion
dollars you're putting down will be you
utilized, would be performant, would be
incredibly cost-effective, and have
useful life for as long as you could see
that infrastructure investment you could
make on Nvidia. You could make with
complete confidence.
We have now proven that it is the only
infrastructure in the world that you
could go anywhere in the world and build
with complete confidence. You want to
put it in any of the clouds, we're
delighted by that. You want to put it on
prem, we're happy about that. You want
to put it in any country, anywhere,
we're delighted to support you. We are
now
a computing platform that runs all of
AI. Now, our business
already starting to show that 60% of our
business is hyperscalers. The top five
hyperscalers.
However, even within that top five
hyperscalers, some of it is internal AI
consumption. The internal AI consumption
really important work like Rexus is
moving from recommener systems of tables
and collaborative filtering and content
filtering. It's moving towards deep
learning and large language models.
Search moving to deep learning large
language models. Almost all of these
different hypers scale workloads are now
moving shifting towards a workload that
Nvidia GPUs are incredibly good at. But
on top of that, because we work with
every AI lab, because we work with every
we accelerate a every AI model and
because we have a large ecosystem of AI
natives that we work with that we can
bring to the clouds that investment no
matter how large, no matter how quick
that compute will be consumed and that
represents 60% of our business. The
other 40% is just everywhere. Regional
clouds, sovereign clouds, enterprise,
industrial, robotics, edge, big systems,
supercomputing systems, small servers,
enterprise servers.
The number of systems, incredible.
The diversity of AI
is also its resilience.
The span of reach of AI is its
resilience. There is no question this is
not a one app technology. This is now
fundamental. This is absolutely a new
computing platform shift. Well, our job
is to continue to advance the technology
and one of the most important things
that I mentioned last year was last year
was our year of inference. We dedicated
everything. We took a giant chance and
reinvented while Hopper was at its prime
and it was just cooking. We decided that
the Hopper architecture the MVL link by
8 had to be taken to the next level. We
completely rearchitected the system,
disagregated the computing system alto
together and created MVLink 72. The way
that it's built, the way it's
manufactured, the way it's programmed
completely changed. Grace Blackwell
MVLink72 was a giant bet and it wasn't
easy for anybody and many of my partners
here in the room. I want to thank all of
you for the hard work that you guys did.
Thank you.
MVLink72
MV FP4 not just FP4 precision FP4 is a
whole different type of tensor core and
computational unit. We've demonstrated
now that we can inference NVFP4
without loss of precision but gigantic
boost in performance and energy
efficiency. We've also been able to use
MVFP4 for training. So MVLink72, MVFP4,
the invention of Dynamo, Tensor RTLM, a
whole bunch of new algorithms. We even
built a supercomputer to help us
optimize kernels and help us optimize
our complete stack. We call it DGX
cloud. We invested billions of dollars
of supercomputing capability help us
create the kernels, the software that
made inference possible. Well,
the results all came together and people
told people used to tell me but Jensen
inference is so easy. Inference is the
ultimate hard. Inference is ultimate
hard. It is also ultimate important
because it drives your revenues. And so
this is the outcome. This is from semi
analysis. This is the largest most
comprehensive sweep of AI that has AI
inference that has ever been done. And
what you see here on the left on on this
side on this side is tokens per watt.
Tokens per watt is important because
every data center every single factory
by definition is power constrained. A
one gawatt factory will never become
two. It's physically constrained the
laws of atoms, the laws of physicality.
And so that one gigawatt of data center
you want to drive the maximum number of
tokens which is the production the
product of that factory. So you want
that you want to be on top of that curve
as high as you want. This the x- axis is
the interactivity the speed of inference
the speed of each inference. The faster
you can inference,
the faster you could of course respond.
But very importantly, the faster you can
inference, the larger the models, the
more context you could process, the more
tokens you can think through. This axis
is the same as smartness of the AI. And
so this is the throughput of the AI.
This is the smartness of the AI. Notice
the smarter the AI, the lower your
throughput. Makes sense? you're thinking
longer. Okay? And so this axis is the
speed. And I'm going to come back to
this. This is important. This is where I
torture all of you. But it's too
important. Every CEO in the world you
watch, every CEO in the world will study
their business from now on in the way
I'm about to describe
because this is your token factory. This
is your AI factory. This is your
revenues. There's no question about that
going forward. And so this is the
throughput. This is the intelligence.
Better perf per watt for a given power
of data center. The more throughput, the
more tokens you could produce. On this
side is cost. Notice Nvidia is the
highest performance in the world. Nobody
would be surprised by that. They would
be surprised by the fact that in one
generation whereas Moore's law would
have given us through transistors 50%
two times
Moore's law would probably give us one
and a half times more performance. You
would have expected from Hopper H200 one
and a half times higher. Nobody would
have expected 35 times higher. I said
last year at this time that Nvidia's
Grace Blackwell NVLink 72 was 35 times
perf per watt. Nobody believed me. And
then semi-analysis came out and Dylan
Patel had a quote.
He accused me of sandbagging.
He accused me of sandbagging. He says,
"Jensen sandbagged. It's actually 50
times." And he's not wrong. He's not
wrong. And so our cost per token, yeah,
our cost per token is the lowest in the
world. You can't beat it.
I've said before, if you have the wrong
architecture, even if it's free, it's
not cheap enough. And the reason for
that is because no matter what happens,
you still have to build a gigawatt data
center. You still have to build build a
gigawatt factory. And that gigawatt
factory for 15 years advertised across
that gigawatt factory is about $40
billion. Even when you put nothing on
it, it's $40 billion in. You better make
for darn sure you put the best computer
system on that thing so that you could
have the best token cost. Nvidia's token
cost is world class.
basically untouchable at the moment. And
the reason that true is because of
extreme code design. And so I'm very
happy that he named us.
There was a monkey king,
token king.
Well, we take we take all of our
software as I as I told you, we
vertically integrate, but we
horizontally open. We're vertical
integration, horizontal open. We
integrate all of our software and all of
our technology, however we could package
it up and integrate it into the world's
inference service providers. And these
these companies are growing so fast.
They're growing so fast. Fireworks. Lynn
is here together. They're just growing
so incredibly fast. A hundred times in
the last year. They are token factories.
And the effectiveness, the performance
and the token cost production capability
for their factories is everything to
them. And this is what happened.
This is we updated their software, same
system.
And notice
their token speeds.
Incredible. The difference before before
Nvidia updated everything and all of our
algorithms and software and all the
technology that we bring to bear
about 700 tokens per second average went
to nearly 5,000 7 times higher. And so
this is the incredible power of extreme
code design. I mentioned earlier the
importance of factories. This is the
importance of factory. Your data center,
it used to be a data center for files.
It's now a factory to generate tokens.
Your factory is limited no matter what.
Everybody's looking for land, power, and
shell. Once you build it, you are power
limited. within that power limited
infrastructure, you better make for darn
sure that your inference because you
know inference is your workload and
tokens is your new commodity that
compute is your revenues that you want
to make sure that the architecture is as
optimized as you can in the future.
every single CSP,
every single computer company, every
single cloud company, every single AI
company, every single
company period are going to be thinking
about their token factory effectiveness.
This is your factory in the future. And
the reason why I know that is because
everybody in this room is powered by
intelligence. And in the future, that
intelligence will be augmented by
tokens. So, let me show you how we got
here.
On April 6th, 2016, a decade ago, we
introduced DGX1,
the world's first computer designed for
deep learning.
Eight Pascal GPUs connected with the
first generation NVLink.
170 teraflops in one computer. The
world's first computer designed for AI
researchers.
With Volulta, we introduced NVLink
switch. 16 GPUs connected with full
alltoall bandwidth operating as one
giant GPU. A giant step forward, but
model sizes continued to grow. The data
center needed to become a single unit of
computing. So, Melanox joined Nvidia.
In 2020, DGXA100 Super Pod became the
first GPU supercomput combining scale up
and scale out architecture.
NVL link 3 for scale up, connect X6 and
Quantum Infiniban for scale out.
Then Hopper, the first GPU with the FP8
Transformer engine that launched the
generative AI era. MVLink 4, Connect X7,
Bluefield 3 DPUs, second generation
quantum infiniband. It revolutionized
computing.
Blackwell redefined AI supercomputing
system architecture with NVLink 72. 72
GPUs connected by NVLink spine 130
terabytes per second of all to all
bandwidth.
Compute trays integrate Blackwell GPUs,
Grace CPUs, Connect X8, and Bluefield 3.
Scale Out runs over Spectrum 4 Ethernet.
With three scaling laws in full steam,
pre-training, post-training, and
inference, and now Agentic systems,
compute demand continues to grow
exponentially.
And now Vera Rubin
architected for every phase of Agentic
AI advancing every pillar of computing
including CPU storage networking and
security.
Vera Rubin Nvlink 72 3.6 exoflops of
compute 260 tab per second of alltoall
NVLink bandwidth the engine
supercharging the era of Agentic AI. The
Vera CPU rack designed for orchestration
and agentic workflows. The STX rack AI
native storage built with Bluefield 4.
Scale out with Spectrum X co-ackaged
optics increasing energy efficiency and
resiliency. And an incredible new
addition, the Gro 3 LPX rack. Tightly
connected to Vera Rubin, Gro's LPU's
massive onchip SRAMM, a token
accelerator to the already incredibly
fast Vera Rubin. Together, 35 times more
throughput per megawatt. The new Vera
Rubin platform. Seven chips, five rack
scale computers, one revolutionary AI
supercomput for agentic AI.
40 million times more compute in just 10
years.
Now, in the in the good old days when I
would say hopper, I would hold up a
chip.
That's just adorable.
This is Vera Rubin. When we think ver
when we when we think Vera Rubin, we
think the entire system vertically
integrated
completely with software
extended end to end optimized as one
giant system. The reason why it's
designed for agentic systems is very
clear because agents of course the most
important workload is it's thinking the
large language model. The large language
models are going to larger and larger
and larger. It's going to generate more
and more tokens more quickly so it could
think more quickly. But it also has to
access memory. It's going to pound on
memory really hard. KV cache structured
data QDF unstructured data QVS. It's
going to be pounding on the me on the
storage system really really hard which
is the reason why we reinvented the
storage system. It is also going to use
tools and unlike humans that are more
tolerant to slower computers.
AI wants the tools to be as fast as
possible. These tools web browsers in
the future they could also be virtual
PCs in the cloud. Those PCs have to be
and those computers have to be as fast
as possible. We created a brand new CPU.
A brand new CPU that's designed for
extremely high singlethreaded
performance,
incredibly
high data output, incredibly good at
data processing, and extreme energy
efficiency. It is the only data center
CPU in the world that uses LPDDR5,
LPDDR5 and incredible single thread
performance and performance per watt
that is unrivaled.
And so that's we built that so that it
could go along with the rest of these
racks for agentic processing. And so
here it is. This is the Grace Blackwell.
Oh no, Vera Rubin. Where is it? Here it
is. Okay, so this is the Vera Rubin
system. Notice since the last time 100%
liquid cooled. All of the cables gone.
What used to take what used to take
2 days to install now takes two hours.
Incredible. And so the manufacturing
cycle time going to dramatically reduce.
This is also a supercomput that is
cooled by it's cooled by hot water 45°
which takes the pressure off of the data
center takes all of that cost and all of
that energy that's used to cool the data
center and makes it available for the
system. This is the secret sauce. It is
the only we're the only company in the
world that has today built the sixth
sixth generation scaleup switching
system. This is not Ethernet. This is
not Infiniban. This is MVLink. This is
the sixth generation MVLink. This is
insanely hard to do. Well, it is
insanely hard to do. Period. And I'm
just super proud of the team. MVLink
completely cooled. This is the brand new
Gro system. And I'll show you a little
bit more about it. this system.
Eight GU chips. This is the LP30. The
world's never seen it. Anything that the
world's ever seen is V1. This is third
generation.
And we're in volume production now. And
I'll show you more about that in just a
second. The world's first
CPO
Spectrum X switch. This is also in full
production. Co-packaged optics. Optics
comes directly onto this chip,
interfaces directly to silicon.
Electrons gets translated to photons and
it gets directly directly connected to
this chip. We invented the process
technology with TSMC. We're the only one
in production with it today. It's called
coupe. It's completely revolutionary.
Nvidia is in full production with
Spectrum X.
This is the Vera system. Twice the
performance per watt of any any CPUs in
the world today. It is also in
production. Well, you know, we never we
never thought we would be selling CPUs
standalone. Um, we are selling a lot of
CPU standalone. This is already for sure
going to be a multi-billion dollar
business for us. So, I'm very very
pleased with our CPU architects. We
designed a revolutionary CPU and this is
the CX9
powered with Vera CPU, the Bluefield 4
STX, our new storage platform. Okay, so
these are the four these are the the the
racks and it's connected
each one of these racks, the MVL link
rack.
This is I've shown you guys this before.
It's a super heavy and seems to get
heavier every year.
because I think there's just more cables
in there every year. And so, so this is
the MVLink rack. We've also taken this
technology because it it is so
efficient to create a data center with
these cabling systems, structured
cables. So, we decided to do that for
Ethernet. So, this is Ethernet, 256
liquid cooled nodes in one rack. And it
is also connected with these incredible
connectors.
You guys want to see um
Reuben Ultra.
So this is the Reuben Ultra compute
node. Unlike Reuben that slides in
horizontally, Ruben Ultra goes into a
whole new rack. It's called Kyber that
enables us to connect 144 GPUs in one
MVLink domain. And so the Kyber rack,
this I I could lift it, I'm sure, but I
won't.
It's quite heavy. This This is one
compute node, and it slides into the
Kyber rack vertically.
This is where it connects into. This is
the midplane. The Kyber racks, those
four top MVLink connectors slide in and
connect into this. And this becomes one
of the nodes.
And each one of these racks is a
different compute node. And this is the
amazing part. This is the midplane.
And the back of the midplane, instead of
the cabling system,
which has its limits in terms of how far
we could drive cables, copper cables, we
now have this system to connect 144
GPUs. This is the new MVLink. This sits
also vertically and it con connects into
the midplanes on the back. Compute in
the front, MVLink switches in the back.
One giant computer. Okay. So that is
Reuben Ultra
as I mentioned. as I mentioned.
How about we t take this back down?
I need the rest of my slides.
>> Oh, it's coming down. Okay. Thank you,
Janine.
This is what happens when you This is
what happens when you don't practice.
Okay. All right. So, um you saw you
Take your time. Just don't get hurt.
You saw you saw this slide. You know,
only at Nvidia's keynote will you see
last year's slide presented again. And
the reason for that is I just want to
let you know that last year I told you
something very, very important. And it's
so important. It's worthwhile to tell
you again.
This is probably the single most
important chart for the future of AI
factories. And every CEO, every CEO in
the world will be tracking it. We'll be
studying it very deeply. It's much much
more complicated than this. It's
multi-dimensional.
But you will be studying the throughput
and this token speed of your AI
factories. The throughput, token speed
at ISO power because that's all the
power you have. Throughput and token
speed for your factories forever. And
that that analysis is going to lead
directly to your revenues. What you do
this year will show up precisely next
year as your revenues. And this chart is
what it's all about. And I said on the
vertical axis, on the vertical axis,
thank you guys. On the vertical axis is
throughput. On the horizontal axis is
token rate. Today I'm going to show you
this
because we're able because we're now
able to increase the token speed and
because model sizes are increasing
because the token length the context
length depending on the different grades
of a different application use case
continues to grow from maybe a 100,000
tokens input length to maybe millions.
the token input length is growing and
also the output token length is growing.
And so all of these play into ultimately
the marketing and the pricing of future
tokens. Tokens are the new commodity and
like all commodities once it reaches an
inflection once it becomes mature or
becomes maturing it will segment into
different parts. The high throughput
low speed could be used for the free
tier. The next tier could be the medium
tier. Larger model, maybe higher speed
for sure, larger input context length.
That translates to a different price
point. You could see from all the
different services, this one is free.
It's a free tier. The first tier could
be $3 per million tokens. The next tier
could be $6 per million tokens. You
would like to be able to keep pushing
this boundary because the larger the
model smarter, the more input token
context length, more relevant, the
higher the speed, the long the more you
can think and iterate smarter AI models.
So this is about smarter AI models. And
when you have smarter AI models, each
one of these clicks allows you to
increase the price. So this is $45. And
maybe one day there'll be a premium
model that allows you a premium service
that allows you to generate token speeds
that are incredibly high because you're
in a critical path or maybe you're doing
really long research and $150 per
million tokens is just not a thing. So
let's translate that. Suppose you were
to use 50 million tokens per day as a
researcher at $150 per million tokens.
As it turns out, as a research team,
that's not even a thing. So, we believe
that this is the future. This is where
AI wants to go. This is where it is
today.
It had to start here to establish the
value and establish it usefulness and
get better and better and better. In the
future, you're going to see most
services encompass encompass all of
that. This is Hopper.
Hopper started and I moved it moved the
chart. This is 50. This is 100. Hopper
looks like this. And you would have
expected Hopper the next generation to
be higher, but nobody would have
expected it to be that much higher. This
is Grace Blackwell. What Grace Blackwell
did is at your free tier increase your
throughput tremendously.
However,
where you mostly monetize your service,
it increased your throughput by 35
times. This is no different than any
product that every company makes. The
higher the tier, the higher the quality,
the higher the performance, the lower
the volume, the lower the capacity. And
so it is no different than any other
business in the world. And so now we're
able to increase this tier by 35x.
And we introduced a whole new tier.
This this is the benefit of Grace
Blackwell. A huge jump over Hopper.
Well, this is what we're doing with
Okay. So, this is Grace Blackwell. Okay.
Let me just reset reset this.
And this is Vera Rubin.
Okay.
Now, just think just think what just
happened at every single tier. At every
single tier, at every single tier, we
increase the throughput. And at the tier
that where your highest ASP and your
most valuable segment, we increased it
by 10x.
That is the hard work. This is
incredibly hard to do out here. This is
the benefit of EVL 72. This is the
benefit of extremely low latency. This
is the benefit of extreme code design
that we could shift the entire area up.
Now, what does it mean from a customer
perspective in the end? Suppose I were
to take all of that and I just, you
know, multiply it against suppose I took
25% of my power, used it in free tier,
25% of my power in the medium tier, 25%
of my power in the high tier and 25% of
my power in the premium tier. My data
center only has a gigawatt.
And so I get to decide how I want to
distribute. The free tier allows me to
attract more customers.
This allows me to serve my most valuable
customers.
And the combination, the product of all
that allows you basically your revenues,
the revenues you can generate, assuming
this simplistic example, allows
Blackwell to generate five times more
revenues.
Vera Rubin to generate five times. Yeah.
So if you're a Reuben, you should get
there as soon as you can. And the reason
for that is because your your cost of
tokens goes down and your throughput
goes up now. But we want even more. We
want even more. And so let me just show
you back to this. This is as you as I as
I told you this throughput requires a
ton of flops. This latency, this
interactivity requires enormous amount
of bandwidth. Computers don't like
extreme amount of flops, extreme amount
of bandwidth because there's only so
much surface area for chips that any
systems has. And so optimizing for high
throughput and optimizing for low
latency are in fact enemies of each
other. And so this is what happened when
we combined with rock. Okay. And so we
we acquired the team that worked on the
Gro chips and licensed the technology
and we've been working together now to
integrate the system. This is what that
looks like. So at the most valuable tier
at the most valuable tier we're now
going to increase performance by 35x.
Now this very simple chart revealed to
you exactly the reason why Nvidia is so
strong in the vast majority of the
workloads so far. And the reason for
that is because up in this area
throughput matters so much. MVLink 72 is
so gamechanging. It is exactly the right
architecture and it's even hard to beat
even as you add Grock to it. However,
if you extended this chart way out here
and you said you wanted to have services
that delivers not 400 tokens per second
but a thousand tokens per second, all of
a sudden MVLink72 runs out of steam and
it simply can't get there. We just don't
have enough bandwidth. And so this is
where Grock comes in and this is what
happens when we push that out. So it
goes out beyond Thank you.
goes out beyond even the limits of what
MVLink72 can do. And if you were to do
that, translate that into revenues
relative to Blackwell Vera Rubin is 5x.
If most of your workload is high
throughput, I would stick with just 100%
Vera Rubin. If a lot of your workload
wants to be coding and very high valued
engineering token generation, I would
add Grock to it. I would add Grock to
maybe 25% of my total data center. The
rest of my data center is all 100% Vera
Rubin. And so that gives you a sense of
how you would add Grock to Vera Rubin
and extend its performance and extend
its value even more. This is what
happens.
Ver this is a contrast. The reason why
the reason why Grock was so attractive
to me is because their computing system
a deterministic data flow processor it
is statically compiled. It is compiler
scheduled meaning the compiler figures
out when the data when to do the compute
the the compute and the data arrives at
the same time. All of that is done
statically in advance
and scheduled completely in software.
There's no dynamic scheduling.
The architecture is designed with
massive amounts of SRAMM. It is designed
just for inference. This one workload.
Now, this one workload, as it turns out,
is the workload of AI factories. And as
the world continues to increase the
amount of high-speed tokens it wants to
generate with super smart tokens it
wants to generate, the value of this
integration is going to get even higher.
And so these are two extreme processors.
You could see one chip 500 megabytes,
one Vera Ruben chip, one Ruben chip 288
gigabytes.
It would take a lot of rock chips to be
able to hold the parameter size of
Reuben as well as all of the context
that has to go the KV cache that has to
go along with it. So that limited
Grock's ability to really reach the
mainstream to really take off until we
had a great idea. What if we
disagregated inference altogether with a
piece of software called Dynamo? What if
we rearchitected the way that inference
is done in the pipeline? so that we
could put the work that makes perfect
sense on Vera Rubin and then offload the
decode generation the low latency the
bandwidth limited challenged part of the
workload for Grock and so we united
unified
two processors of extreme differences
one for high throughput one for low
latency it still doesn't change the fact
that we need a lot of memory and so
Grock we're just going
add a whole bunch of Grock chips which
expands the amount of memory it has and
so if you could just imagine
out of a trillion parameter model we
have to store all of that in gro chips
however it sits next to Nvidia Vera
Rubin where we could we could hold the
massive amounts of KV cache that's
necessary in processing all of these
agentic AI systems it's based upon this
idea of this aggregated inference we do
the prefill that's the easy part but we
also tightly integrate the decode so the
attention part of decode is done on
Nvidia's Vera Rubin which needs a lot of
math and the feed forward network part
of it the decode part is done uh the
token generation part is done on Vera
Rubin on the uh on the groip the two of
them working tightly coupled together
over today Ethernet with a special mode
to reduce its latency by about half. And
so that capability allows us to
integrate these two systems. We run
Dynamo, this incredible operating system
for AI factories on top of it. And you
get 35 times increase. 35 times
increase. Not to mention additional new
tiers of inference performance for token
generation the world's never seen. So
this is it. This is Grock.
the Vera Rubin systems including Grock.
I want to thank Samsung uh who
manufactures the Gro LP30 chip for us
and they're cranking as hard as they
can. I really appreciate appreciate you
guys. We're in production with the Gro
chip and uh you know we'll ship it in
the second half probably about Q3 time
frame. Okay.
Grock LPX
Vera Rubin you know it's kind of hard
it's kind of hard to imagine any more
customers
you know and and uh the the really great
thing is is um Grace Blackwell early
sampling of it was really complicated
because of coming together of Envy Link
72 but the sampling of Vera Rubin is
just going incredibly well and in fact
Satia I think texted out already that
the first Vera Rubin rack is already up
and running at Microsoft Azure and so
I'm super excited for them. We're just
going to keep cranking these things out.
We have now set up a supply chain that
could manufacture thousands a week of
these systems essentially multi-
gigawatts of AI factories per month
inside our supply chain. And so we're
going to crank out these these Vera
Rubin racks while we're cranking out the
GB300 racks. We are in full production.
The Vera CPUs
incredibly successful. And the reason
for that is because AI needs CPUs for
tool use and Vera CPU was designed just
perfectly for that sweet spot.
Incredible for the next generation of
data processing. Vera CPU is ideal. the
Vera CPU plus blue plus CX9 connected
into the Bluefield fourstack
100%
100% of the world's storage industry is
joining us on this system and the reason
for that is because they see exactly the
same thing. The storage system is going
to get pounded. It's going to get
pounded because we used to have humans
using the storage systems. We used to
have humans using SQL. Now we're going
to have AIS using these storage systems
and it's going to store QDF accelerated
storage, QVS accelerated storage as well
as very importantly KV caching. Okay, so
this is the Vera Rubin system. Now
what's amazing is this. in just two
years time in a one gigawatt factory in
just two years time in one gigawatt
factory using the mathematics that I
showed you earlier whereas Moore's law
would have given us a couple of steps we
would have you know x factored the
number of transistors we would have x
factored the number of flops we would
have x factored the number of amount of
bandwidth but with this architecture
we're going to take our token generation
ation speed token generation rate from 2
million to 700 million 350 times
increase.
This is this is the power of extreme
code design. This is what I mean when we
integrate and optimize vertically but
then we open it horizontally for
everybody to enjoy. This is our road
map. Very quickly,
Blackwell is here, the Oberon system. In
the case of Reuben, we have the Oberon
system. We're always backwards
compatible. So that if you wanted to not
change anything and just keep on moving
through with the new architecture, you
could do so.
The old the standard um rack system
Oberon still available. Oberon is copper
scale up. And with Oberon, we could also
use optical scale out or excuse me,
optical scale up to expand to MVLink
576.
Okay. And so there's a lot of
conversation about is Nvidia going to
copper scale up or optical scale up.
We're going to do both.
So, we're going to have MVLink 144 with
Kyber and then with Operon uh Opteron
Oberon, we're going to MVLink72
plus Optical to get to MVLink 576.
The next generation of Reuben with
Reuben Ultra. We have the Reuben Ultra
chip which is coming which is imp taping
out and we have a brand new chip LP35.
LP35 will for the first time incorporate
Nvidia's MVFP4 computing structure give
you another few X X factor speed up.
Okay. And so this is Oberon MVLink 72
optical scale up and it uses Spectrum 6
the world's first co-ackaged optical and
um all of this is in production. The
next generation from here
is Fman. Fineman has a new GPU of
course. It also has a new LPU
LP40.
Big step up. Incredible. Incredible new
technology. Now
uniting the scale of Nvidia and the Gro
team building together LP40. It's going
to be incredible. a brand new CPU called
Rosa,
short for Roslin. Bluefield 5, which
connects the next CPU with the next
Superneck
CX10.
We will have Kyber
which is copper scale up. We will also
have Kyber
CPO scale up. So for the first time we
will scale up with both copper and
co-ackage optics. Okay. And so a lot of
people have been asking you know Jensen
are is copper going to still be
important? The answer is yes.
Jensen are you going to scale up
optical? Yes.
Are you going to scale out optical?
Yes.
And so for everybody who is in our
ecosystem, we need a lot more capacity
and that's really the key. We need a lot
more capacity for cop for copper. We
need a lot more capacity for optics. We
need a lot more capacity for CPO and
that's the reason why we've been working
with all of you to lay the foundation
for this level of growth. And so Fman
will have all of that. Let me see if I
uh missed everything. That's it. every
single year. Brand new architecture.
Very quick.
Very quickly, Nvidia went from a chip
company to a AI factory company or AI
infrastructure company, AI computing
company. These systems
and now we're building entire AI
factories. There's so much power
that is squandered in these AI
factories. We want to make sure that
these AI factories come together
designed in the best possible way. Most
of these components never meet each
other. Most of most of us technology
vendors now we all know each other but
in the past we never met each other
until the data center. That can't
happen. We're building super complex
systems and so we have to meet each
other virtually somewhere else and so we
created Omniverse and the Omniverse DSX
world a platform where all of us can
meet and design these gigafactories the
giga you know gigawatt AI factories
virtually in system. We have simulation
systems for the racks for mechanical,
thermal, electrical, networking. Those
simulation systems integrated into all
of our ecosystem partners of incredible
tools companies. We also operated,
connected to the grid so that we could
interact with each other, send each
other information so that we could
adjust
grid power and data center power
accordingly, saving energy. And then
inside the data center using Max Q so
that we could adjust the system
dynamically across power and cooling and
all of the different technologies we all
work on together so that we leave no
power squandered
so that we run at the most optimal rate
to deliver enormous amount of token
throughput. There's no question in my
mind there's a factor of two in here and
a factor of two at the scale we're
talking about is gigantic. We call this
the NVIDIA DXX platform. And just as all
of our platforms, there's the hardware
layer, there's the library layer, and
there's the ecosystem layer. It's
exactly the same way. Let's show it to
you.
The greatest infrastructure buildout in
history is underway.
The world is racing to build chip system
and AI factories. And every month of
delay costs billions in lost revenues.
AI factory revenues are equal to tokens
per watt. So with power constraints,
every unused watt is revenue lost.
NVIDIA DSX is an Omniverse digital twin
blueprint for designing and operating AI
factories for maximum token throughput,
resilience, and energy efficiency.
Developers connect through several APIs.
DSX SIM for physical, electrical,
thermal and network simulation, DSX
exchange for AI factory operational
data, DSX Flex for secure dynamic power
management between the grid and DSX Max
Q to dynamically maximize token
throughput.
It starts with SIM ready assets from
NVIDIA and equipment manufacturers
managed by PTC windshield PLM.
Then modelbased systems engineering is
done in DASO systems 3D experience.
Jacobs brings the data into their custom
Omniverse app to finalize design.
It's tested with leading simulation
tools using Seaman's Star CCCM Plus for
external thermals,
Cadence Reality for internal,
EAP for electrical, and NVIDIA's network
simulator DSX Air
and virtually commission through Procore
to ensure accelerated construction time.
When the site goes live, the digital
twin becomes the operator. AI agents
work with DSX Max Q to dynamically
orchestrate infrastructure.
Fedra's agent overseas cooling and
electrical systems, sending signals to
Max Q, which continuously optimizes
compute throughput and energy
efficiency.
Emerald AI agents interpret live grid
demand and stress signals and adjust
power dynamically.
With DSX, Nvidia and our ecosystem of
partners are racing to build AI
infrastructure around the world,
ensuring extreme resiliency, efficiency,
and throughput.
It's incredible, right? Well,
om Omniverse Omniverse was designed to
hold the world's digital twin starting
from the earth and it's going to hold
digital twins of all sizes. And so we
have just such a great ecosystem of
partners. I want to thank all of you.
All of these companies are brand new to
our world. We didn't know many of you
just a couple years ago. And now we're
working so close together to work on and
build together the largest computer the
world's ever seen and also to do it at
planetary scale. So NVIDIA DSX is our
new AI factory platform.
I'll spend very little time on this at
this time. However, we're going to
space. We've already been out in space.
uh Thor is radiation approved and uh
we're in satellites. You do imaging from
the from satellites. In the future,
we'll also build data centers in the in
the in the in space. Uh obviously very
complicated to do. So we have we're
working with our partners on a new
computer called Vera Rubin Space 1 and
it's going to go out to space and start
data centers out out in space. Now, of
course, in space, there's no conduction,
there's no convection, there's just
radiation. And so, we have to figure out
how to um uh cool these systems uh out
in space. But, we've got lots of great
engineers working on it. Let me talk to
you about something new.
So, so um
uh Peter Steinberger is here and um uh
he he wrote a piece of software. It's
called Open Claw and and um I don't know
if he realized uh how successful it was
going to be. Um but the importance is
profound. Open Claw is the number one.
It's the most popular opensource project
in the history of humanity and it did so
in just a few weeks.
It exceeded it exceeded what Linux did
in 30 years. And it's that important. It
is that important. It will do well.
Uh this is all you do. Okay? We're
announcing our support of it. Uh let me
let me just quickly go through this.
this. I want to show you a couple
things. You simply type this, you type
it this into a into a console and um it
goes out, it finds open claw, it
downloads it, it builds you an AI agent,
and then you could tell it whatever else
you need to do. Okay, so let's take a
look.
An open source project just dropped.
>> Andre Carpathy has just launched
something called research is a huge
deal.
>> You give an AI agent a task, go to
sleep, it runs 100 experiments
overnight, keeping what works and
killing what doesn't.
I really love what my stuff enables that
person to do. And he had like one guy,
he told me like he installed it as a
60-year-old dad and like they made beer,
connected the machine via Bluetooth to
open claw. And then we automated
everything including the whole website
for people to order lobster.
Hundreds of people are queuing up for
lobsters in s openclaw.
>> Open claw.
>> You want to build open claw with open
claw.
>> Everyone is talking about open claw. But
what is open claw?
>> Believe it or not, there's already a
claw con.
Incredible. Incredible. Now, um I
illustrated effectively what open claw
is in this way and so all of you can
understand it. But let's just think what
happened. What is open claw? It connects
it's an a it's a system. It calls and
connects to large language models. So
the first thing it has it has resources
that it manages. It manage it could
access tools. It could access file
systems. It could access large language
models. It It's able to do scheduling.
It's able to do cron jobs. It's able to
um uh decompose a problem that a prompt
that you gave it into step by step by
step. It could spawn off and call upon
other sub aents.
It has IO. You could talk to it in any
modality you want. You could wave at it
and it understands you. You could talk
to any modality you want. It sends you
messages, it texts you, sends you email.
So, it's got IO.
Um, what else does it have? Well, based
on that, you could you could say in fact
it's an operating system. I've just used
the same syntax that I would describe an
operating system. Art
openclaw has open sourced essentially
the operating system of agent computers.
It is no different than how Windows made
it possible for us to create personal
computers. Now open claw has made it
possible for us to create personal
agents.
The implication is incredible. The
implication is incredible. First of all,
the adoption says something you know all
in itself. However, the most important
thing is this. Every single company now
realize every single company, every
single software company, every single
technology company for the CEOs, the
question is what's your open claw
strategy?
Just as we need to all have a Linux
strategy, we all needed to have a HTTP
HTML strategy which started the
internet. We all needed to have a
Kubernetes strategy which made it
possible for mobile cloud to happen.
Every company in the world today needs
to have an open claw strategy and a
gentic system strategy. This is the new
computer. Now this is just the exciting
part. This is enterprise IT before
openclaw you know and and I mentioned
earlier the way enterprise IT works and
the the reason these reason why it's
called data centers is because these
large rooms these large buildings held
data held the files of people the
structured data of business. It would
pass through software that has tools and
you know systems of records and all
kinds of workflow that's codified into
it and that turns into tools that humans
would use
digital workers would use. That is the
old IT industry software companies
creating tools saving files and of
course gsis consultants that help
companies figure out how to use these
tools and integrate these tools. These
in these tools are incredibly valuable
for governance and security and privacy
and compliance and all of that's
continues to be true.
It's just that post open clock post
agentic this is what it's going to look
like. This is the extraordinary part.
Every single IT company, every single
company, every SAS company,
every SAS company will become a
a gas company.
No question about it. Every single SAS
company will become a gas company, an
agentic as a service company. And what's
amazing is this. You now open claw gave
us gave the industry exactly what it
needed at exactly the time.
Just as Linux gave the industry exactly
what it needed at exactly the time just
as Kubernetes showed up at exactly the
right time just as HTML showed up it
made it possible for the entire industry
to grab onto this open-source stack and
go do something with it. There's just
one catch.
Agentic systems
in the corporate network can have access
to sensitive information. It can execute
code and it can communicate externally.
Just say that out loud. Okay, think
about it. Access sensitive information,
execute code, communicate externally.
You could of course access employee
information,
access supply chain, access finance
information, sensitive information and
send it out, communicate externally.
Obviously,
this can't possibly be allowed. And so,
what we did was we worked with Peter. We
took some of the world's best security
and computing experts and we worked with
Peter to make open claw
open claw enterprise
secure and enterprise private capable.
And we call that
this is our Nvidia open claw reference
for open nemo claw which is a reference
for openclaw and it has all these
agentic AI toolkits and the first part
of it is technology we call open shell
that has now been integrated into open
claw now it's enterprise ready this
stack this stack with a reference design
we call Nemo cloud neoclaw
Okay, with a reference stack we call
Nemo clock. You could download it, play
with it, and you could connect to it the
policy engine of all of the SAS
companies in the world. And your policy
engines are super important, super
valuable. So the policy engines could be
connected Nemo Claw or Open Claw with
Open Shell would be able to execute that
policy engine. It has a polic
guard rail. It has a privacy router and
as a result we could protect and keep
the the clause from executing inside our
company and do it safely. We also added
several things to the agent system and
one of the most important things you
want to do with your own
claw custom claws is so that you can
have your custom models and this is
Nvidia's open model initiative. We are
now at the frontier of every single
domain of AI models. Whether it's
Neimotron, Cosmos, World Foundation
model, Groot, artificial general
robotics, human or robotics models,
Alpamo for autonomous vehicle, Bioneo
for digital biology,
Earth 2 for AI physics. We are at the
frontier on every single one. Take a
look.
The world is diverse. No single model
can serve every industry.
Open Models is one of the largest and
most diverse AI ecosystems in the world.
Nearly 3 million open models across
language, vision, biology, physics, and
autonomous systems enable AI builds for
specialized domains. NVIDIA is one of
the largest contributors to open-source
AI. We build and release six families of
open frontier models, plus the training
data, recipes, and frameworks to help
developers customize and adopt new
leaderboard topping models are launching
for every family. At the core, Neotron
reasoning models for language, visual
understanding, rag,
safety,
and speech.
>> Can you hear me now? Hello. Yes, I can
hear you now.
>> Cosmos Frontier models for physical AI
world generation and understanding.
Alpayo, the world's first thinking and
reasoning autonomous vehicle AI
group foundation models for general
purpose robots. Bioneo open models for
biology, chemistry, and molecular
design.
Earth 2 models for weather and climate
forecasting rooted in AI physics.
NVIDIA open models give researchers and
developers the foundation to build and
deploy AI for their own specialized
domains.
Our models our mo thank you
our models are valuable to all of you
because number one it's on the top of
the leaderboard. It's world class. But
most importantly, it's because we are
not going to give up working on it.
We're going to keep on working on it
every single day. Neotron 3 is going to
be followed by Neotron 4. Cosmos one was
followed by Cosmos 2. Groot Groot at
generation 2. Each and one of these
we're going to continue to advance these
models. vertical integration,
horizontal openness, so that we can
enable everybody to join the AI
revolution, number one on leaderboard
across research and voice and world
models and artificial general robotics
and self-driving cars and reasoning and
of course one of the most important one.
This is Neotron 3 in
Open Claw. This is Neimotron 3 and Open
Claw. And look at the top three. There
are the three best models in the world.
Okay. So, we are at the frontier.
It is also true. It is also true that we
want to create the foundation model so
that all of you could fine-tune it and
post-train it into exactly the
intelligence you need. This is Neotron 3
Ultra. It is going to be the best base
model the world's ever created. This
allows us to help every country build
their sovereign AI. And we're working
with so many different companies out
there. And one of the most exciting
things that we're doing today, I'm
announcing today is a Neotron coalition.
We are so dedicated to this. We have
invested billions of dollars of AI
infrastructure so that we could develop
the core engines for AI that's necessary
for all the libraries of inference and
so on. But also to create the AI models
to activate every single industry in the
world. Large language models is really
important. Of course, it's important.
How could how could human intelligence
not be? However, in different industries
around the world, in different countries
around the world, you need to have the
ability to customize your own models and
the domains of the domain of the domain
of the models is radically different
from biology to physics to self-driving
cars to general robotics to of course
human language. And we have the ability
to work with every single region to
create their domain specific their
sovereign AI. Today we're announcing a
coalition to partner with us to make
Neotron 4 even more amazing. And that
coalition has some amazing companies in
it. Black Forest Labs imaging company.
Cursor the famous coding company we use
lots of it. Lang chain billion downloads
for creating custom agents. Mistrol the
Arthur Arthur mentioned I think he's
here. Incredible incredible company.
Perplexity Perplexes computer absolutely
use it everybody use it. It is so good.
A multimodal agentic system. Reflection
Sarv from India thinking machine mirror
Morardi's lab. Incredible companies
joining us. Thank you.
I said I said that every single
enterprise company, every single
software company in the world needs an
agentic systems, need an agent strategy.
you need to have an open claw strategy.
And they all agree
and they're all partnering with us to
integrate Nemo, the Nemo claw reference
design, the NVIDIA agentic AI toolkit,
and of course all of our open models.
One company after another. There's so
many. And we're partnering with all of
you. I'm really grateful for that. And
um this is our moment. This is a
reinvention. This is this is a
renaissance
a renaissance of the enterprise IT from
what would be a $2 trillion industry.
This is going to become a multi-
trillion dollar industry offering not
just tools for people to use but agents
that are specialized in very special
domains that you're expert in that we
could rent. I could totally imagine in
the future every single engineer in our
company will need an annual token
budget.
They're going to make a few hundred,000
a year their base pay. I'm going to give
them probably half of that on top of it
as tokens so that they could be
amplified 10x. Of course, we would. It
is now one of the recruiting tools in
Silicon Valley. how many tokens comes
along with my job. And the reason for
that is very clear because every
engineer that has access to tokens will
be more productive and those tokens as
you know will be produced by AI
factories that all of you and us we
partner to build. Okay. So every single
enterprise company in today sit on top
of file systems and data centers. Every
single software company of the future
will be agentic and they will be token
manufacturers. They'll be token users
for their engineers and they'll be token
manufacturers for all of their
customers. The open clause in event, the
open claw event cannot be understated.
This is as big of a deal as HTML. This
is as big of a deal as Linux. We have
now a world-class open agentic framework
that all of us could use to build our
open claw strategy. And we've created a
reference design we call Nemo cloud
neoclaw that all of you could use that
is optimized. It's performant. It is
safe and secure.
Speaking of agents, agents as you know
perceive, reason and act. Most of the
agents in the world today that I've
spoken about are digital agents. They
act in the digital world. They reason.
They write software. It's all digital.
But we also have been working on
physically embodied agents for a long
time. We call them robots. And the AIs
that they need are physical AIs. We have
some big announcements here. I'm going
to just walk through a few of them. 110
robots here. Almost every single company
in the world, I can't think of one that
are building robots is working with
Nvidia. We have three computers. The
training computer, the synthetic data
generation and simulation computer, and
of course the robotics computer that
sits inside the robot itself. We have
all the software stacks necessary to do
so. the AI models to help you.
And all of this is integrated into
ecosystems around the world and all of
our partners from Seammens to Cadence,
incredible partners everywhere. And
today, we're announcing a whole bunch of
new new partners. As you know, we've
been working on self-driving cars for a
long time. The Chad GPT moment of
self-driving cars has arrived. We now
know we could successfully autonomously
drive cars. And today we are announcing
four new partners for Nvidia's robo taxi
ready platform.
BYD,
Hyundai,
Nissan,
Ji all together, 18 million cars built
each year. joining our partners from
before Mercedes, Toyota,
GM. The number of robo taxi ready cars
in the future are going to be
incredible. And we're announcing also a
big partnership with Uber.
Multiple cities were going to be
deploying and connecting these robo taxi
ready vehicles into their network. And
so a whole bunch of new cars. We have uh
ABB, Universal Robotics, uh CUKA, so
many robotics companies here and we're
working with them to implement our
physical AI models integrated into
simulation system so that we could
deploy these robots into manufacturing
lines all over. We have Caterpillar
here. We even have T-Mobile here. And
the reason for that is in the future
that radio radio tower used to be a
radio tower is going to be an NVIDIA
aerial AI ram. And so this is going to
be a robotics radio tower. Meaning it
can reason about the traffic, figures
out how to adjust its beam forming so
that it could save as much energy as
possible and increase the amount of
fidelity as possible. There's so many
humanoid robots here, but one of my
favorites, one of my favorites is a
Disney robot. You know what? Tell you
what, let me just show you some of the
videos. Let's look at that first.
The first global rollout of physical AI
at scale is here. Autonomous vehicles.
And with NVIDIA Alpamo, vehicles now
have reasoning, helping them operate
safely and intelligently across
scenarios.
We ask the car to narrate its actions.
>> I'm changing lanes to the right to
follow my route.
>> Explain its thinking as it makes
decisions.
>> There's a double parked vehicle in my
lane. I'm going around it.
>> And follow instructions.
>> Hey, Mercedes. Can you speed up?
>> Sure, I'll speed up.
>> This is the age of physical AI and
robotics.
Around the world, developers are
building robots of every kind. But the
real world is massively diverse,
unpredictable, full of edge cases. Real
world data will never be enough to train
for every scenario.
We need data generated from AI and
simulation. For robots, compute is data.
Developers pre-trained World Foundation
models on internet scale video and human
demonstrations and evaluate the model's
performance to prepare them for
post-training.
Using classical and neural simulation,
they generate massive amounts of
synthetic data and train policies at
scale.
To accelerate developers, Nvidia built
open-source Isaac lab for robot training
and evaluation and simulation.
Newton for extensible and GPU
accelerated differentiable physics
simulation.
Cosmos world models for neural
simulation
and Groot open robotics foundation
models for robot reasoning and action
generation.
With enough compute, developers
everywhere are closing the physical AI
data gap.
Paratas AI trains their operating room
assistant robot in NVIDIA Isaac Lab,
multiplying their data with NVIDIA
Cosmos World models. Skilled AI uses
Isaac Lab and Cosmos to generate
post-training data for their skilled AI
brain. They use reinforcement learning
to harden the model across thousands of
variations. Humanoid
uses Isaac Lab to train whole body
control and manipulation policies.
Hexagon Robotics uses Isaac Lab for
training and data generation.
Foxcon fine-tunes group models in Isaac
Lab,
as does Noble Machines.
Disney research uses their chamino
physics simulator in Newton and Isaac
lab to train policies across their
character robots in every universe.
Da
Does
>> ladies and gentlemen Olaf
>> does coming through Newton Newton works.
>> Wow.
>> Omniverse works.
Olaf,
how are you?
>> I'm so happy now that I'm meeting you.
>> I know because I gave you your computer,
Jetson.
>> What's that?
Well, it's in your tummy.
>> That's going to be amazing.
>> And you learn how to walk inside
Omniverse.
>> I love walking. This is so much better
than riding on a reindeer gazing up at a
beautiful sky.
And
it was because of physics using this
Newton solver that runs on top of Nvidia
Warp that we jointly developed with
Disney and with DeepMind that made it
possible for you to be able to adapt to
the physical world. Check that out.
>> Not to say that
that's how smart you are.
>> I'm a snowman, not a snowed.
Could you imagine this? The future of
Disneyland. All these all these robots,
all these characters wandering around.
>> Oh,
>> you know, I have to admit though, I
thought you were going to be taller.
I've never seen such a short snowman, to
be honest.
>> Nope.
>> Hey, tell you what. You want to help me
out?
>> Hooray.
>> Okay. Usually usually I close the
keynote by talk telling you what I told
you. We talked about inference and
flection. We talked about the AI
factory. We talked about the open claw
agent revolution that's happening. And
of course we talked about physical AI
and robotics. But tell you what, why
don't we get some friends to help us
close it out?
>> Of course.
>> All right, play it.
Come on.
>> Terminating simulation.
Hello.
Anybody here?
The keynotes over all was said. Jensen
map the road ahead. AI factories coming
alive. Agents learning how to drive from
open models to robots too. Now we'll
break it all down.
Comput exploded. What we saw from CNN's
to open cloth agents working across the
land but they need the power to meet
demand. So we saw the problem it was
brilliant. We multiplied compute by 40
million.
But once upon AI time training was the
paradigm. Sure it talk models how in
France runs the whole world now shows us
who's the bars at 35 times less the cost
blackwell makes the token singing video
the inference king.
Yeah, our factories once took years as
vendors pulling racks and gears. Built
up slowly, piece by piece. No clear way
to scale this beast. DSX and Dynamo know
what to do.
Turning power
into revenue.
Agents used to wait and see, now act
autonomously. But if they ever try to
stray, safe claws block and say no way.
Nemo claws there to guard the course.
And yes, my friends,
it's open sorrow.
Cars that think and droids that run.
This ain't the movies. It's all begun.
Alamo calls the shots. It's a GPT moment
for the bots from sim streets. Now watch
them drive. Blow your hands up
for physical AI.
Industrial age. Build what came before.
Now we build for AI. Even more vera
rubin plus grog make the inference
splash put them together now it's
raining cash we build new architecture
every year cuz claws keep yelling more
tokens here the AI stacks for all to
make so let us all eat five layer cake
the moment's bright the path is clear
cuz open models led us here when data's
missing there's no dispute we just
generate more with compute robots
learning without flaw fueling the four
scaling laws the future's here won't you
come and see welcome Welcome all to GTC.
All right, have a great GTC
wave.
Thank you everybody.
I just met
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.