Building the Cloud for AI Agents | AWS CEO Matt Garman
Before automating a workflow with AI, redesign it from the outcome backward. Pick one recurring task, write the desired result and constraints, then ask how an agent could explore several paths in parallel rather than merely copying a person’s sequence of steps. Start with narrow, time-boxed permiss
56mSummary published by 1% Better, updated .
Key Takeaway
Before automating a workflow with AI, redesign it from the outcome backward. Pick one recurring task, write the desired result and constraints, then ask how an agent could explore several paths in parallel rather than merely copying a person’s sequence of steps. Start with narrow, time-boxed permissions, a sandbox, and a human review point. This creates useful speed without granting an untested agent broad access to production systems or sensitive data.
Episode Overview
AWS CEO Matt Garman discusses how AI agents are changing cloud infrastructure, software development, enterprise operations, and the economics of compute. He explains AWS’s approach to GPU allocation, custom chips, data-center constraints, agent security, and the organizational implications of teams increasingly managing fleets of coding agents.
Main Insights
Ranked strongest first for usefulness, specificity, and support in the episode.
1. Redesign workflows for parallel agents, not human imitation
Garman says enterprises often begin by asking an agent to replicate Bob’s steps one through five, with Bob reviewing the result at the end. He argues the higher-value approach is to start from the desired outcome and let an agent try many approaches in parallel, which changes how a computer can solve the problem rather than simply digitizing an existing process.
2. Give agents narrow, temporary permissions
Agent permissions should not simply inherit a person’s full access, according to Garman. He describes time-boxed, task-specific, fine-grained permissions and sandboxing as new building blocks for letting agents act without exposing an organization to unnecessary risk.
3. Treat evaluation as a production capability
For autonomous enterprise agents, Garman highlights continuous testing, labeled data, production measurement, back-testing, and monitoring for drift. The implication is that an agent deployment is not finished at launch: teams need an ongoing feedback loop that tests whether real-world behavior still matches the intended goal.
4. Start simple, then grow into governance
AWS is reducing initial setup friction by allowing new accounts to start with defaults, while preserving the ability to add organizational controls and fine-tune configuration later. This is a useful product-design principle: remove early barriers without forcing users into a separate, disposable path that requires migration when they scale.
5. Build for transient work without sacrificing a production path
Agents may create a database for a short task and discard it, while production systems require durability and availability. Garman frames the design challenge as supporting rapid, lightweight experimentation while allowing the same underlying environment to become a durable production system if the work proves valuable.
6. Use AI to accelerate security, not only to assess risk
Garman acknowledges that powerful models can expand the attack surface, but says they can also help identify and prioritize vulnerabilities. His example is using AI with context about permissions, architecture, and compensating controls so security teams can focus on the vulnerabilities that matter most.
7. Keep proprietary data inside a trusted boundary
Garman argues that enterprise data is a company’s most valuable asset and says customers should be deliberate about where prompts and data travel. AWS’s stated Bedrock approach keeps customer data in the customer’s VPC rather than returning it to the model provider, illustrating the governance questions teams should resolve before moving from prototypes to production.
8. Consider customized open-weight models only when you can validate them
Garman sees promise in post-training or fine-tuning open-weight models on proprietary data to potentially improve performance and lower cost. But he explicitly says the organization needs a good evaluation system to prove that result, making measurement a prerequisite rather than assuming customization automatically delivers value.
9. Use agents to unblock domain experts
Inside Amazon, Garman sees HR and finance teams building agents for work that previously required software-development support, such as team planning, resource management, and tax-rule compliance. The practical opportunity is to equip knowledgeable business users with safe, governed tools so they can solve problems directly instead of waiting in a development queue.
10. Organize smaller teams around faster-moving problems
Garman says agentic development can allow three or four people to build what once required a ten-person product team, creating a need to move people more fluidly across projects. He does not claim to have a final organizational model, but advises experimenting with structures that balance maintaining existing systems with rapidly pursuing new opportunities.
Notable Quotes
"You want to actually step back and say, if I want to accomplish something, how can an agent do it differently? It can do it in a massively parallelized way. It can try 50 different things and get to that."
"You don't just want to give it Raghu's permissions and let it go do whatever you can do. You actually want, you know, very time box short-term permissions to just go do a task."
"Customers are going to need security at machine speed, not at human speed, not at like an alarm, someone goes in, look at it."
"It's not code completion. It really is agent first, the agents write all of the code. You're just managing a team of agents and driving that."
Action Items
-
1
Redesign one workflow from the outcome backward
Choose a repetitive process this week. Define the desired output, constraints, failure modes, and review criteria; then list ways an agent could perform research, comparison, or drafting in parallel instead of copying the current human sequence.
-
2
Create an agent access policy
For each agent, document the single task it may perform, the systems and data it may access, the minimum permissions required, an expiration time, and actions that require human approval. Run it first in a sandboxed environment.
-
3
Build a lightweight evaluation loop
Collect representative examples of good and bad outcomes, define pass/fail criteria, test the agent before release, and sample production outputs regularly. Track failures and drift, then update prompts, tools, permissions, or data accordingly.
-
4
Find one domain-expert bottleneck
Ask an HR, finance, operations, or customer-support colleague to identify a task delayed by engineering capacity. Pilot a constrained agent that helps retrieve, summarize, classify, or check information while keeping a person responsible for final decisions.
Full Transcript
Transcript of Building the Cloud for AI Agents | AWS CEO Matt Garman from A16Z. Auto-generated from episode audio; may contain minor errors.
Agentic workflows tend to perform better on AWS than anywhere else. Compute sandboxes, gateways, agent permissions versus people. A lot of those are things that we have built and are building and thinking actively about. The top frontier labs gobble up every available GPU. At the same time, you want to promote newer companies that are going to become the enterprises of tomorrow. How are you thinking about balancing? From the very beginning of when we launched AWS, startups have been the lifeblood of the core of what we do.
These are the innovators that are at the edge of technology, understanding what's possible. We're very intentional about keeping capacity available for the startups. We recently announced we're going to be buying 2 million NVIDIA GPUs over the next couple of years. Your cap ex this year is what, 200 or something? $220 billion for 26. And we don't anticipate slowing down anytime soon because the demand is just massive. With all the debate around AI, extension risks and the hugging face attack, what are CEOs asking you about all these things?
Amazon plans to spend $220 billion in capital this year. And AWS CEO Matt Garvin says they don't anticipate slowing down anytime soon. In this episode, v16z's Raghu Raghuram sits down with Matt to unpack what's driving one of the largest infrastructure buildouts in history and how AI is changing the cloud itself. They discuss how AWS decides who gets scarce GPU capacity. They're shifting bottlenecks from power and chips to memory and construction. And the company's bet on custom silicon electroneum. They also get into what changes when agents increasingly write code and manage infrastructure.
And that shares what AWS is seeing inside its own teams, where agents now write code and engineers increasingly manage teams of agents, accelerating how quickly new products can be built. Welcome to the pod, Matt. What a time we're living in. So I have a lot of topics to talk to you about. Awesome. Thanks for having me. I'm excited. Yeah, absolutely. So let's start actually right from the beginning. So you are the first GM for EC2. And that was 2006, right? And today you guys are what, 160, 170 billion in revenue?
Yeah, about 169, 170 billion. Yeah, growing 35, 37%. That's 37% at 169. Crazy. It's interesting to think about it from day one when we got the first dollar of revenue. But yeah, the interesting thing is we're still at the early stages of what the business can be and what the opportunity is for our customers. Most workloads, there's a huge amount of workloads still that live on-prem today. And the amount of compute that people are doing every single day is more than it was the day before.
And so you see the tailwind from AI, you see the tailwind from migration into the cloud, and the business has grown really fast. And it's been a super fun thing to be a part of. Yeah, it is. I mean, they'll be writing history books about this and business books about this for a long time to come. I want to touch on the on-prem. That sort of boggles my mind. I obviously did my best to keep them there for a long time. You built a lot of stuff on-prem back in the day.
Yeah, yeah. Trying to now get all of that to move into AWS. Yeah, so we can talk about that later. So if you think about EC2 in the early days, and you got your start with obviously selling to startups. And today, the AWS cloud service, I mean, as you reflect on the evolution, what stands out? Part one, and then part two, we'll talk about how this nature of how you serve startups has changed. Sure. Well, like you said, actually, it's funny. So I actually interned for AWS in 2005, when we were first kind of was an internal project.
It was my business school internship. And my project was actually to come up with an analysis of who we thought AWS would be most interesting to. And the answer was startups. Surprisingly. And so from the very beginning of when we launched AWS, startups have been the lifeblood of the core of what we do for a number of reasons. One is the value proposition is just so attractive, what AWS provides to startups. And we spend a lot of time and effort making sure that we're great partners to the startups, helping them not just provide infrastructure, but also advice on how to get your company up and running and how to think about your architecture so it'll scale eventually.
And a whole bunch of things that we do for the startups. We also think that for us, it's just good business because the startups today are the enterprises of tomorrow. And it's an imprecise number. But today, we estimate that maybe 30-40% of AWS revenue comes from companies that was a one-time startup in AWS's lifetime. So it's fun to see for us to see the companies grow over time and to get bigger and become enterprises effectively. And so that's why we invest so much in startups and why we pay so much attention to the brand new two people in a garage type startups.
And not just because of business outcomes, but also because that's who we learn from. These are the innovators that are at the edge of technology, understanding what's possible, pushing our services to say what they'd want more of, what can help them go faster, what can help them achieve their outcomes. And a lot of times, they're pushing more than the banks or the healthcare companies or the governments. The startups are the ones that are pushing that envelope, and it really helps us to be better and make sure that we're ahead of that wave where the banks are going to want some capability five years down the road that startups want today.
Yeah, yeah. And compared from then to now, how have the startups changed and what they want from you? Obviously, they all want a lot of GPUs, and we can talk about that, but besides that. Well, I will say there's a couple of things that have changed. One is startups started at a much smaller size than when we first started, right? They might have gotten $10 million of funding, and they had an app idea that they were kind of slowly iterating on. Now, it's from day one, they're valued at a billion dollars.
They have $200 million of funding. It's a team and an idea, and all of a sudden, they're worth a billion dollars. Let's go from our offices to your offices. That's right. That's right. And the size and scale and ambition of the ideas, I think, requires more capital. They're bigger to start with. They're obviously more expensive, too, right, to go after kind of training a model or doing something that a lot of the folks are doing today. So that's number one, which is just the size that they start from is really big.
I would say number two is some things haven't changed, right? They actually are still thinking about, okay, once I scale, how do I think about an architecture? How do I think about security? How do I think about performance? How do I think about having all the capabilities that I need? How do I think about setting up my IAM setup so that when I have more than three employees, this thing is going to work and scale, which is a lot of times why they like AWS as opposed to just going to a neocloud or something like that.
They need all of those other security and capabilities that come around from training. And so I think that's something that hasn't changed and is exactly the same. I think the scale at which the ramp up is definitely different today. I also think increasingly we're seeing teams that want a cloud that is great to work with agents and not just with people. And I think that's where we've spent a lot of time. How do you think about, exactly, how do you think about broad scale? How do you think about performance?
How do you think about an interface that is a well-defined API interface that agents can actually easily traverse and work across? And so that's also something that we both, I think, were naturally set up to do well, but that we've actually doubled down on to ensure things like, how can you start a database in three seconds and really get to some capabilities that agents are excited about? Have you introduced any specific new services that are explicitly targeted at agents? Are people building agents? I'd say what we've done is we've optimized some existing services so that they can both work for people and agents.
And so the answer is, we definitely have some services designed for building agents. So we have things like Agent Core and Bedrock and pieces like that. But when we look at just taking an underlying component like S3, which is where the vast majority of companies store their data and have their data lakes, it turns out a lot of the use cases for people and agents are similar. But you want to have, and one of the things that we have in preview right now, or in beta, is called AWS Context.
That allows you to build a context layer so that agents can actually more easily find all of the data they want across your various data lakes. And so whether you have your data stored in Aurora or you have it in S3 or you have it in somewhere else in AWS, you can kind of build this context layer where when people aren't necessarily going to access data in that way, but agents actually are happy to go across lots of those different things. And so there are some services that we're building like that, but for the most part, you also care about what does latency look like?
What does throughput look like? How do you make sure that the underlying engine is actually fast, scalable? They actually care a lot about tail latencies, which is interesting. People don't always care about the P999 S3 latency. Agents do care and get blocked by that. And so that's the thing that we've cared about for a long time. And from a performance perspective, it's one of the things that really popped. That's why agentic workflows tend to perform better on AWS than anywhere else. Obviously, we see a lot of companies starting out, and the very common refrain, of course, is that agents are writing all the code for them, right?
And agents are selecting the databases and the email servers and everything that you can name it, right? So has there been a lot of thinking that – and obviously, they did the documentation. Has there been a lot of thinking on the AWS, like your feels funny to say, quote, unquote, legacy services that have been around for a decade on how do you re-architect them? So, yes, I think there are some things that we've thought about. Actually, the core underlying building blocks, I think, are in really good shape.
I think there's a usability layer that we're thinking about of how we make easier to use for – we'll call it the very simplest use case you call out where somebody is just coding up an app really quickly and they ask it to deploy. We find the vast majority of our customers will go and tell their coding agent, whether it's Kiro, whether it's Claude, whether it's Codex, they'll say, I want to build on AWS. Here's my credentials. Here's my stuff. This is how I want to deploy it.
And then agents are great, and they go do it. There are use cases, though, where you're already not on a cloud. You're just trying something, and you say deploy. And frankly, a lot of times that will go to some of our partners that have an easier to use layer on top, which we love, by the way, too. And we love our partners on those fronts, too. But I do think that there are some things that we're doing where if you're brand new, you don't already have an AWS, you already haven't set up your IAM.
A couple of months ago, it was much harder to start up an account, right? It's been like this for 20 years. But you have to define your VPCs, you have to define your IAM roles, all these kind of things, which are actually super important things once you get to be large. And what we hear from customers is like it was a hard tradeoff because they know that they're going to want those in months or years down the line. But for now, they just kind of want to use the services, not worry about that.
And so what we've actually launched is if you go create a new AWS account today, and we're slowly rolling this out. I don't know if it's actually fully all the way out yet now. But you don't have to do those things. You don't have to give it a credit card. You can sign up with your Gmail account. All of those things are handled in default behind the scenes. And within less than 30 seconds, you're up and running and can be operating in a full AWS account, which is for those type of systems, much more what the agents want to be able to do because they don't want to have to go through all that setting up your VPCs and pieces like that.
So there are some things where we're adding kind of some ease of use to that, which we're quite excited about. We've seen some really positive feedback from customers on it, and we'll keep doing more things like that over time. The great part about that, though, is it's not like a simplistic account and then you have to migrate. If you basically say, okay, now I actually do want to scale into the org, I want to develop an organization, I want to go and kind of fine tune some of the things, you can easily come just, you're already in a real AWS account, and you can actually then go and do all of those things later when you need them.
And so there's no migration or move later. And so that's part of the hard work that we really think about, which is how do we make it not a choice for the customer, but an easy on-ramp into the depth of features that we know customers and startups are going to want when they start scaling. Yeah, yeah, got it. And what has been some of the hardest things to accommodate as agents have taken over, I mean, you made a name, the way you guys became, the phenomenon that you became was serving developers.
And then infrateams, now developers are substituted by agents, and pretty soon infrateams are being substituted by agents. Are there, what has been some of the hardest things for you as the largest service provider in the world to handle that transition? Well, look, I think, like I said, much of our infrastructure is actually pretty well set up to handle the scale, which is good. I think there's been a couple of things which are interesting, which are, I think, interesting paradigm shifts, where you could argue that some of our systems are, I wouldn't say over-engineered, but as an example, many agents want to create a database, do a little bit of work, and then have the database go away.
You really need that database to have five nines of durability. Exactly. Right, and so there's some things that we rethink there where when you're creating an Aurora database, like as your production database, you do want five nines of data. You need durability, you want availability, you want all of those things. And for the agent use case, that's arguably, or maybe not arguably, over-engineered for what we need. And so it's hard, we don't, and we kind of have this belief that we don't really want to have a non-durable option that's going to cause problems either, because you never really quite know if the database that's created wants to stay around for a long time or a short amount of time.
And so we're trying to do the hard work to think about how do you accomplish both of those things where you can create it quickly and throw it away and you don't really waste a lot of resources. But if you do want it and it can be durable and stay for a long time and actually grow into a big production database. So those are some trade-offs that we think about actively as we think about how does the more traditional kind of, this is going to be my production system, kind of capabilities match with some of the more transient nature of the infrastructure that agents want to use.
So that's one, I think, but there's a bunch. I think the scale and speed and latency of creation of durable resources is another one that's interesting. You know, I think the other one that we actively think about too is just are there new building blocks that agents are going to want that we just didn't really need before. And so compute sandboxes, gateways, agent permissions versus people or service role permissions, a lot of those are things that we have built and are building and thinking actively about because they are just brand new building blocks.
It's not using the existing building blocks differently, but it's brand new ones where pretty clearly you want a different permissions for agents. You don't just want to give it Raghu's permissions and let it go do whatever you can do. You actually want, you know, very time box short-term permissions to just go do a task. You actually may not want to give permissions to a whole tool at all. You may want to get very fine-grained commissions for what an agent's able to do from a sandbox, right?
You actually want a sandbox that you want to... The organization is having its act too. That's right, exactly. And they're different. It's similar. It's the same idea, but it's, you know, how do you have that lightweight and think about... Unfortunately, we've done a lot of work with Firecracker, our kind of micro VMs. A lot of the sandbox companies, startups use Firecracker. They all use Firecracker, which we invented 10 years ago, maybe, or something like that. And it's really purpose-built for... I mean, it wasn't purpose-built for agents, but it's actually quite good for agents because you can spin them up really rapidly, they have a great security boundary, and you don't have a lot of virtualization overhead that you have from a traditional large VM.
So, you know, I think those are some new building blocks that we're thinking about and more emerge every single day. Yeah. So, I've rattled enough into agents. I'll come back to it later, but it's such a fascinating topic. Super cool. But let me switch gears a little bit and ask about something that every one of our startups face, and we get asked most frequently, which is, how can we get GPUs, right? Yeah. I mean, obviously, you guys are running massive GPU farms, increasing it every day, but the top frontier labs gobble up every available GPUs, more power to them.
How do you deal with the internal... at the same time, you want to promote newer companies that are going to become the enterprises of tomorrow, like you said. Yeah. It's a great question. Well, yeah. It's a great question. And a couple of these things make it more challenging too. One is the CapEx expense needed to go deploy kind of all of the compute that everyone needs right now is massive, right? And so we are... I think... 220 billion dollars for 26. It's... At one point, I saw that it's a pretty...
I mean, that is a larger expense than we've ever had, maybe any company has ever had in a single year. And we don't anticipate slowing down anytime soon because the demand is just massive. And so at some point, you're limited by how fast can you build data centers, how fast can you deploy capital, how fast can you get memory and chips and all of those kind of things. And so all of those are different points, constraints that we're... whether it's core data centers, capital, memory, chips, they're all kind of constraints at various times.
Construction people that build buildings like construction people are at a premium today. And so all of those, we work really hard to make sure happen. And then with the capacity that we are able to deploy, which is still a huge amount and not enough, we know, we really think intentionally about what is that allocation strategy. And so we're great partners with the large Frontier Labs, the Anthropics and Open Source and Meta and other large customers. And so those folks are really good customers of ours. And we want to make sure that we invest in them.
We have large enterprise customers, whether it's Salesforce's or JPMC's or other large companies that have demand for fewer numbers of GPUs or accelerators. Sometimes they want Tranium, sometimes they want NVIDIA GPUs. But we want to make sure that we can support them as well. And we're very intentional about how we make sure we have the capacity for startups. And so what we do is we actually do allocate and we basically say, OK, we're going to keep, you're right, we could sell every single GPU or AI accelerator we had to probably just the big Frontier Labs and call it a day.
We choose not to do that because we actually want to keep growing the full ecosystem. They get a large number, but we want to keep supporting a broader set of customers because we actually think that both the whole ecosystem will be more healthy for us. There's some diversification, but it's also just we know these are going to be big companies over time. And so we try to support them. I saw recently that we say, yes, in some way, shape, or form to something like 60% of the requests we eventually get.
Sometimes it's a little bit later. Sometimes it's in a different region. Sometimes it's a slightly different configuration than the customers are looking for. But we really try to lean in and try to allocate as much as we can. And every single startup that you have wants more. Exactly. It's been hard to try to make sure that we have all of those, but it's hard. But this is, as a lot of people say, it's a good problem to have. Yes, it is. But it's a problem nonetheless.
And we're we continue to look at we recently announced we're going to be buying two million NVIDIA GPUs over the next couple of years. We're landing a massive amount of capacity. Two million. So it's a lot. And it's over the next couple of years. And who knows if that's enough. At some point, again, we're limited by other components as well. But it is. We're very intentional about keeping capacity available for the startups. And I know it's painful to not have enough, but we keep pushing. When did your CapEx cross I mean, pure AWS.
I mean, Amazon was a bigger company. Crossed, let's call it 1 billion or 10 billion a year. Oh, I don't know. I'd have to go back and look. I'm not sure about that. But it's definitely scaled up over the last couple of years in a pretty meaningful way. The AI build-out has definitely ramped our CapEx spending. And so I'd say over the last call, three or four years, our CapEx has definitely accelerated pretty meaningfully. I don't know when we crossed a billion. But given we're at 220 now, it's probably a little while.
I mean, we've been spending CapEx for a while. And the company has been really passionate about funding that. And obviously, the AWS business is a good one that we like to invest in. Yeah, of course. And I think we were, for the longest time, actually, we were investing ahead of where the demand was. And I think one of the most painful things is that with the real ramp of GPUs, a lot of the elasticity has unfortunately gone away. And so hopefully we'll get back. And in our core compute and performance, that elasticity is still there.
But we've been spending for quite a bit of time. And we feel really good about the spend that we're making now. Which is, I think I get lots of questions sometimes about how you feel about that spend. And like, are you nervous about bubble and other things like that? And I will say, because of the position we have, one, we do take this diversified approach. And so not all of our capacity is bundled up in one customer. And I think some of these, whether they're NeoClouds or some of the other providers out there, and you'll see sometimes concentrations of 30%, 40%, 50%, 60% with one or two customers.
We're nowhere near that, obviously. We're single digit percentages at the highest and usually it's less than that. For one, I think we have a lot less risk on one particular customer. But also, because we have that rich set of services, AWS is where people really are coming to launch their production workloads. And so the majority of our usage today, actually, is either core compute and storage and inference, which is part of that application. And so those are the workloads that I think just aren't going to go away.
Because we see enterprises getting positive ROI. You go talk to the customers and you say, at the capability today and the cost today, are you seeing positive returns to your business? And almost to a person, they'll say, like, oh yeah. And so you're like, well, that's not going to go away. There's no bubble in which they stop spending on that. And you know this, the VC model, like, is every billion dollar startup going to make it? No, they won't. But, you know, that's kind of the game.
That's been true for 50 years. They haven't always been billion dollar startups. The numbers have changed, but the principles haven't changed. But the principles are the same, right? You bet on 10 and one makes it and pays for the other 10 or whatever the percentage is, hopefully it's higher than 1 out of 10. Yeah. Right. And so you saw that with the internet where there was a bubble and a bunch of internet companies didn't make it. And the internet is still a thing. And a lot of the companies that had durable businesses, the Google, the Amazons, the others, like, they did pretty well.
And so for us, we think we'll feel really good about that investment and the continued investment going forward. You guys have a view of the demand that's unparalleled, right? Because you're seeing across the globe, you're seeing across every segment, enterprise and the big labs and the native companies and so on and so forth. So if anybody should call it, you should be able to call it first. I would hope so. That's our plan. And honestly, we spend a lot of time thinking about it. We're very intentional about how we spend our shareholder capital.
And we think we're making great investments. We have a lot of good protections of how we intentionally invest that money. But yeah, we're very bullish. And I think Andy's been public about saying this. The potential for AWS is really, really large. And over the next decade, the potential is there. And anytime you have an opportunity that's that big, you want to invest to go after it. As the scale of these numbers go larger and larger, has your planning and process dramatically changed in terms of – I mean, you're not writing billion-dollar checks.
You're writing like $50 billion checks or $20 billion checks or whatever it is. Yes and no. I mean, I think, look, a lot of times we're still very bullish about the investments and lean forward. But the process – so there's a lot of things that have completely changed. How do you even estimate 2028 demand? And there's things that we have to think about now that we just never had to think about. And so if you go back 15 years, if we needed more power, we asked the power company to give us another couple of megawatts or whatever it was, right?
Tens of megawatts. And they would just give it to us because 10 megawatts wasn't that much. And that was plenty for us to keep growing. Now we have to bring our own. We bring our own power. And so we pay for power projects. We pay for renewable projects. We're one of the biggest renewable power purchasers each year for the last 10 years. And so it's – we're regularly bringing on new solar projects, new nuclear projects, new – So are you talking about behind the meter or are you talking about working with the operator?
We'll do both. And so both of those things. And so oftentimes with these power projects that we bring on, we'll pay the capital and pay for the project, and then it'll go into the grid, and then we'll get credit for that. So we'll bring those on. Sometimes we'll do behind the meter too. It's a mix. At the scale that we're doing, you have to think about all of those things. But, you know, that's planning where you're thinking 20 years out of how you're going to think about power, how you're going to think about transmission, how you're going to think about that capital, and that's stuff we've never had to think about before.
So that's planning we just never had to do. But, you know, we always had to – the other thing is, when we used to think about server demand that we needed, we'd have multiple quarter demand things, and we'd talk with our suppliers and things. Now we work multiple years out just because the size is so much that we have to think about kind of what do we need for 26, what do we need for 27, what do we need for 28? But it's also one of the values that we bring to customers.
That's a thing that legitimately customers can't do themselves. They're not going to do power planning. They're not going to plan their memory footprint in 2028. They can't do that. So that's one of the real values that we bring to our broad set of customers. That's just a whole set of things that you don't have to worry about and that we spend a huge amount of time thinking about. Yeah, yeah. You guys are generally, for the record, the largest buyers of practically every component of a server, correct?
I don't know that. I am sure that we're one of the bigger purchasers of components out there, for sure. And who knows about everyone? And that kind of depends on how you think about them and how you measure. So what I was leading to is where do you see the constraints being most severe in, let's call it, 27, 28, and where do you see the constraints easing up? Yeah, it's funny. So I'll answer this a roundabout way, but I remember that in undergrad, we actually, this was a long time ago, and I never thought this would be a useful book that I read.
But we read the goal. Oh, of course. And so it turns out there's never one constraint. There's always just the latest constraint. And so you have to think about all of them. And as soon as you hit one, there's another one, right? And so what is the constraint that's going to happen in 20 and 28? I actually don't think there will be one. I think it's like every month for us, it's do you have enough power? And then as soon as power is no longer the constraint, it might be memory.
It might be TSMC capacity. It might be HBM. It might be networking components. It could be there's a blip somewhere in the supply chain and connectors or whatever it is. At some point, you have to think about all of those pieces. And it's not also where they happen matters, too. It may be like, hey, we have a ton of power in Indonesia, but we don't have enough in Germany. And so you think about where in the world you want that capacity, too, because it turns out that not everything is totally fungible.
Some is, and some is not. But it might be disk drives. It might be SSDs. We think about every single component, and we have tens, hundreds of thousands of components, all that we track and think about. And some we rely on our suppliers to manage. Much of it, we directly manage. And yeah, we have a whole team that does that, and they're fantastic. I think they're industry-leading. And we saw this problem coming probably a decade ago and really started not just thinking about, okay, how many servers do we need to track, but just thinking all the way through the supply chain, four tiers, five tiers down.
What is the component that could cause an issue for us, and making sure that we had guaranteed supply on that. If you think, if you remember, gosh, I don't remember when this was, it was over a decade ago when there was the floods in Thailand and no one had disk drives anymore. There was disk drive crisis, and then there was memory crisis. Exactly. And so I think we think through all of those things, and so we also think about, where is there diversification in manufacturing? All of those kind of pieces we try to work through, and we're never going to be perfect on it, but there's always a different supply constraint.
So obviously there's a lot of wide-ranging debate about data centers, right? It's clear that folks like us, where we stand, but do you think as an industry, we have not done a good job of explaining why data centers are good for America, generally the world, but and what's the internal talk? amongst Andy's team on how do we deal with this? Yeah, well, look, I think and I think you'll hear more from us over this. And I agree. I think we need to be more vocal and be more upfront because we actually do a ton that's that's really beneficial, both for communities we operate in for the you know, we think a ton about how do we bring renewable energy to these data centers?
How do we think about being water positive? Actually, our data centers use a really, really small amount of water. We mostly use free air cooling. How do we think about being great participants in the communities where we are and and how we bring high paying jobs to the communities we operate in? And not all data center operators do that. Yeah, I think there are some well chronicled examples of others out there that are not great at that and they just don't really pay attention to regulations.
They think that the rules don't apply or they just launch really quickly without thinking about those. And I think that it causes a problem for the whole industry because everybody kind of gets lumped into that. So so, look, I think we'll be we I think you're right. We, you know, vocally self-critical. We need to be more vocal about the benefits that we do bring and think about additional ways that we can help communities understand the benefits that that we bring to them, both for the services they use, right?
If you usually go to a community and say, well, do you not want to use Netflix? And they'll be like, no, no, like I still want Netflix. And, you know, like it's important for us to think about and highlight the benefits that we bring, where I recently saw a report where one of the communities that we operate in. The everybody in that county pays five thousand dollars a year less in taxes because of the taxes that we bring to that, and we don't we don't tell them.
They don't even know it. Yeah, right. It's just invisible to them. And so I think we just need to be more clear about those benefits that we bring, because I think if you told the communities, by the way, your tax bill is five thousand dollars less than it would otherwise be. If we weren't here, they might have a little bit of a different thought about the building that's over there. Yeah, yeah. So and but but not everyone does that. Not everyone kind of. Presumably there's more transparency.
That's what I think much of the data center community, not just us, is actually pretty good actors. And there's just a few that aren't that that that kind of I think have have caused some of the angst recently. And I think we just need to be a better job of highlighting, you know, who's who's being good citizens and who's not. Yeah. And before we leave the hardware topic, I want to touch on Trinium and your whole history with building your own chips. We were one of your first partners using Nitro a long time ago.
And since then, Graviton made tremendous progress. So what was the thinking that led to saying, look, we're going to do our own thing. And then how has that progress been and where are you? Yeah, it's it's actually a fascinating story. And I think it's a great example of where Amazon AWS will innovate and will iterate over time and continue to think bigger about what we can do. But but but kind of prove our way there as opposed to, you know. And so take this as an example.
It's probably now 10, it's probably about 13, 14 years ago. We were seeing that there was a pretty significant virtualization tax on the overall number of resources. And and we were kind of thinking about how do we our customers were telling us, I want bare metal performance. And they, you know, they're comparing having all the resources of a server. And so the first thing that we did is that we took a network offload card and virtualized all of our network virtualization and pulled it off into an offload card so that network virtualization got closer to to bare metal performance.
And back then it was not quite bare metal, but it was close, it was closer. And then we got really excited about that. And we said, OK, what if we could move storage virtualization off as well? Right. And and and none of the network offload cards could do that. And then we we found this one company who had some arm cores on an offload card and they were doing it for other reasons. I can't remember their original purpose. But we're like, could you use those to do storage virtualization and some of these other functions?
And we're like, maybe. And so we really iterate with them. This is the Annapurna team and and just love that team, like really innovative, really mission driven, really wanting to solve problems. And so we acquired them and we said, look, could you build us a slightly bigger card that actually could take all the network virtualization off and basically give us a bare metal server that has no virtualization on it? No VM virtualization. Everything is through APIs on the card. And because we had this view that one performance would be much better.
Research, a resource utilization would be better. The security isolation. And security isolation would be much, much better. And we tell people, you know, could then legitimately tell people we have no access to any of your VMs that are running there. And and and this has been a huge benefit for us for the last decade, where, frankly, like we have been leading and others have been kind of slow to do this because this is not a generalized thing that people can do. But so we got to that and we basically said, look, we're we're making a lot of progress here.
What if we take, you know, and there's a bunch of ARM cores that run this offload card. And we said, what if we turn that into a server? And we did that first with Graviton. It was a very underpowered, very small server that we launched. And customers are excited. They're like, I'd love to have an ARM server. This is super interesting. And so we went down the path and and Graviton, you know, and part of what we did is we looked and saw that there's the the slope of, you know, ARM cores ARM cores were getting faster and where that you saw the power utilization and the graph where you knew the intercept was going to happen for where this architecture was going to be really good for for parts.
And they just needed somebody to drive the ecosystem and get some of the pieces in place. So we did that with Graviton. And Graviton has been a runaway hit at this point. Have you been public about what percent of your fleet is Graviton? We land more Graviton ships every year than than than any other type. So it's it's it's very popular. And it's, you know, look, we they're 20 percent cheaper at a 20 percent better performance and have been like that for the last kind of five, six years.
So that's a it's an easy value proposition that and I think the vast majority, something like 90 plus percent of our top hundred customers all use Graviton in some way, shape or form across their fleets. And so that's been a huge win for us and for customers. It's been the single big single easiest way that customers lower their bill is to move to Graviton. They can offer we've we've had examples where people have moved their whole fleets and cut the number of servers they had in half.
Performance is so much better. It's amazing. And so half his number of servers, each server costs less. Like it's it's a it's a big win. And so then about five, six years ago, we said, you know, we saw the rise of of AI compute happening, not nearly expecting what it was today, but still saw it was going to be a big mover. And so we went in and built our first chip in Tranium, and we're now in market with the third generation in Tranium three. And I've seen fantastic results.
So it's you know, we're we're sold out for capacity through probably towards the end of next year or something like that. And we're trying to again, we're trying to figure out how we can save some capacity and get startups to be able to use some of the capacity because it's we see great results. So most of, you know, the the majority of of traffic on Bedrock all runs on Tranium. And we have great deals with both Anthropic and OpenAI to build on top of Tranium, as well as a set of of smaller startups.
And, you know, I think we have half a dozen to a dozen startups that are building on top of Tranium now, too. And so now the name suggests it's a training chip. Yeah, we're back. Everybody's using it for inference, too. So where is the lean architecturally and where is it going? It's a good point. Look, the vocally self-critical, we're terrible at naming. And so it's not it's not our strength. Originally, we had a chip called Inferentia for inference and training for Tranium for training. And then as the models got bigger and bigger and it turns out you actually want to run the inference on these really large systems.
It turns out that training turns out that Tranium is actually maybe the best inference chip on the market right now from an absolute performance and cost performance point of view. And so is it it's it's the better memory memory bandwidth. Where is the it has it's just the architecture is a little bit different than than than others. And it's much cheaper. And so from a cost performance perspective and absolute performance perspective, Tranium is great. And so, you know, we use it a ton for for inference.
And like I said, Bedrock, it drives much of the Bedrock inference today. And but it's also a good training chip, too. And it's you know, I think we're we we you know, a lot of our broad set of customers usage is not in training models, but is in using it. And so that's where a lot of people get to use under the covers. And that's where we're excited about it. But but a lot of the big customers are interested in it for training clusters as well.
And particularly as you get to Tranium three and four, which we've announced, we haven't launched Tranium four yet, but announced it. Folks have kind of looked at that architecture and said, yeah, that's the future of where my training clusters to be as well. So we're quite excited about the future, where that goes for these really broad scale training clusters also. But yeah, it's both. Now, let's get back to talking about agents, but from a perspective of large enterprises or large medium enterprises. Yeah. Where are they in their adoption and have they what sort of benefits are you seeing them reap already and what is the roadmap for them?
As far as you can tell from your vantage point. Yeah, it's a really good question. I think it's one that. That we've spent a lot of time thinking about, and when I talk to customers all out there today, they view, you know, they're getting a lot of value out of what they've done today, and I would say the agents that most enterprises have built are are relatively simple and straightforward, and they're starting to think about and they're mostly non-autonomous, right? They're still kind of people in the loop, if you will.
And so I think we're and by the way, there's customers are still getting lots of value out of that today. And so they're really thinking about how do I have these be autonomous, but in a safe way? And I think there's two things that I think hold customers back today from just continuing to scale. And it's already a pretty big business today, but I think it has a massive opportunity to really change every single customer out there and every single workflow and really thinking about it.
And so number one is just how to think about it. I think what we originally saw was that enterprises had a workflow and they're saying, great, I would have the enterprise. I would have an agent go do the same workflow. Yeah. And what we encourage them to really do is think not just replicate. You know, Bob does step one, two, three, four or five. So agent is going to do one step one, two, three, four or five. And then Bob's going to check it at the end.
That's not really the model you want. You want to actually step back and say, if I want to accomplish something, how can an agent do it differently? It can do it in a massively parallelized way. It can try 50 different things and get to that. And and how do you help it get to that right outcome and rethink how a computer would solve a problem versus a human solving a problem? And so one of the things is us just helping customers understand how to think about that and really kind of have that blank slate, because that's where you really get value is not just replicating what you're doing today, but but thinking from a Greenfield approach about how you go solve a problem completely differently.
And I'm sure that's how many of your startups are thinking about this, too. How do you help customers Greenfield solve a problem, not replicate the thing that happens today? That is number one. People running fleets of agents and swans, whatever you want to call them. Yeah, very common these days. And you just want to think about it. You know, and so enterprises are not as as again, this is one where you learn from the startups and you try to how do you apply that to an enterprise world where they're, you know, an insurance company is not necessarily as forward leaning, but they would love to figure out how they can have a better, you know, approval workflow or something like that.
So that's number one. But then the second one is this this how do you turn those into fully autonomous workflows and how do you actually trust the agents? And so we're spending a lot of time thinking about how do we build services to help enterprises feel like their systems are secure and that they can trust an agent to make a decision that can have the right guardrails, that can have the right permissions on their data, that is not going to delete production systems, that it's not going to make tragic mistakes.
And right now, I think that nervousness is probably holding people back. That maybe appropriately, by the way, is holding enterprises back from just saying, OK, go nuts like you don't actually want an agent to just go crazy and accidentally delete a production database. That's going to be pretty bad. And so that's how we're kind of actively working through this with customers on how do we both help them architect and frankly, invent new technologies and capabilities that are going to help them solve that problem. And so that is one of the areas where I think we'll continue to innovate and we'll get there.
I think we have some really good ideas and some good technologies brewing that I think can really help. So are enterprises learning how to do eval systems and so on and so forth to keep the agents under control? Honestly, like both like evals, how do you have a constant loop of testing? How do you how do you think about, you know, goal seeking in a reasonable way? How do you have your data labeled in such a way that it actually even makes sense that the eval can actually kind of approximate what you're going to be doing in production?
How do you measure in production and back test it so you're not seeing drift? All of those things are problems that that enterprises don't have to solve today. I don't know if anyone really is great at solving these today. It's why you've seen so many FTE teams kind of spin up. And AWS and our partners are really leaning into the FTE motion to go and help. And this is the single biggest area where customers need help. And when we think about how you and my view is we want to train our customers to be able to go and do this themselves.
Right. This is not the traditional motion where I want to have a people driven business that goes on forever, where you just keep paying consultants over and over and over again. Our view on how FTEs should work and is we want to. Go into a customer who's ready to accept the you know, to really kind of accept owning this when we're done and in 45 days do work where we can teach them how to make an eval, teach them how to get their data in a labeled way, do it, work alongside with them.
And then at the end of 45 days, you leave and that customer is is good and ready to go and trained up. And that's what our customers tell us they want. They don't want to be beholden to an external workforce for the next five years. And but they need help today. Yeah, you've made a massive investment in FTEs. Yeah. So taking that even one step further. Right. Some of your industry peers have said, look, you can't have all of your data going into a big frontier model.
What enterprises should really do is to take an open source model. and then post-train on your own data and workflows and traces and whatnot. Where do you stand on that? Are you seeing customers actually trying to do that? Or do you guys, how do you think about that? It's great. The first point, I wholeheartedly agree on that first point, like the customers and enterprise data is their most valuable asset. And so from the very beginning, it's why we built Bedrock like we did. We have a guarantee that your data never leaves your VPC.
And so if you're running inside of Bedrock, your data doesn't go back to the model provider. They never see your prompts that stays inside of your own trusted environment. And so that is why enterprises kind of run, they prefer to run on top of Bedrock. And it's why you see that business growing massively, like hundreds and hundreds. I mean, it's every month we see that just business, just every week we see that business exploding. And it's why you see OpenAI workloads migrating to Bedrock. It's why you see Anthropic really growing really rapidly.
And so whether you're using open models or closed frontier models, I think Bedrock is a great solution that our customers told us, by the way, like, you know, if you remember three years ago, I got a lot of heat. Speaking of bad names, that's a good name, though. Yeah, Bedrock is a good name. That's good. But we got a lot of heat actually for being slow to the AI world because we actually built the foundations of this where we said, look, we're not just going to rush out of service.
We really want to think about how do we make sure that we protect our customers' data and build a service that we think is going to be durable for the use cases that we knew about. And if you remember, we got a lot of heat and we said, look, we're going to go build the right thing. And now as people move from proof of concepts to production, vast majority of them are landing in AWS on Bedrock. For one of the reasons is because of this. It's also because of the set of services that we have.
We also offer open models. We offer proprietary models. We offer a whole set of capabilities around those. Agent core, we build these building blocks, so it's easier to build agents with any of the models that you want. Whether they're in Bedrock or out of Bedrock, for that matter, you can use Gemini or other things for it. But I think it's a differentiating piece for us and it's a super important thing to think about because having that data go back into the model provider, I think is a dangerous thing.
You talk about open weights models, though. I do think that there's a scenario that I'm excited about where we're really ramping up our support of open weights models and trying to build a good environment. And frankly, this is where today, I think, and I think it's true, a lot of customers believe that they have meaningful proprietary data, that if they could mix in, do some post-training, do some fine-tuning to an open weights model that they could distill down, they can actually get a better performing model at a lower price.
There's a bunch of pieces here that have to work out well. They actually have a good eval to actually prove that that's true. Most people are doing that in SageMaker today. I think there's more that we can do to make that easier. But actually, if you go look at where people are doing that, they actually do it in SageMaker on AWS today. They actually host the inference via SageMaker. SageMaker is getting a new lease of life. I mean, it is. It was always kind of a model building platform.
And now, if you think about what enterprises are doing in this world, that's what they're doing is they're basically effectively building their own custom models. And so SageMaker is a great place of doing that. And I think there's some things that we need to keep building on to make that easier and easier to do and to test across different open weights models, things like that. But it's an evolving space. I think it's a super interesting one, and it's one we want to make sure that we have all the right things for customers to be able to do if they have the right data and expertise to actually go down that path.
And with all the debate around AI, extension risks, and this, that, and the other, and the security vulnerabilities, and so on and so forth, and the hugging face attack. How are enterprises, what are CEOs asking you about all these things? Yeah, there's a bunch. And they look, they mostly want to say like, and it goes back to this, like when I launch agents, how can I trust that they're going to do what I want them to do? And so we're increasingly, we have been for a while, like we're basically, we're heavily investing in building capabilities that allow people to deploy agents safely into their environment and think about those controls.
And some of those are, how do you make sure the agents have the right permissions? How do they have the right sandboxing? How do you make sure that you have the right set of guardrails? How do you really intentionally think about what you want the agents to do and not do? Is there a human in the loop or not? And so we spend a lot of time with our customers thinking about how do you think about safe agent deployment and get better over time? And what other capabilities do we need to go build to help people deploy agents safely into their environment?
So there's a lot there. The other angle on that, which a lot of people are worried about, which is just, are some of these really powerful models going to be attack services and kind of mythosphere? And so I kind of have a view on there. Yes, that is a real risk, I think, to customer environments, but it's also a real opportunity. And so we recently launched a service called Continuum that uses these powerful models to help customers go secure their environment. And so we'll look across their environment, look for vulnerabilities with them.
We'll use some of these powerful models and help customers find vulnerabilities they haven't found before. And most importantly, by the way, prioritize which ones, because we know context about their environment, how it's set up, where their permissions are, where they may have compensating controls that make it harder or easier to do. And so Continuum is incredibly popular with customers. Actually, we're really bullish about what's possible from AI to help with AI powered security, because look, at some point, customers are going to need security at machine speed, not at human speed, not at like an alarm, someone goes in, look at it.
And so that, you know, we're running fast to go help build that for customers to help them protect their environments. And I'm very excited about what the Continuum team is building on that front too. So within AWS itself, like how, what's the state of usage of agents in the browser? Well, Continuum is basically us trying to expose what we do internally. And so we use AI extensively for our own security. We use AI extensively for our own software development. We use agents. Actually, one of the things that's really cool is we use agents across our entire business.
And so we rolled out Amazon Quick to every single Amazon employee. And now I see HR teams building agents to help drive what used to take teams of people weeks to do that a single person can now do in a couple of hours to think about kind of team planning and resource management. I have finance teams that are building agents to go think about how do they go pull tax rules from everywhere and ensure that we have compliance on a bunch of different pieces. And super cool to see that things that used to be blocked by software developers, actually the line of business folks are able to go and unblock themselves and innovate more quickly.
And so Quick has been an enormous blower. And that has grown like wildfire. We see customers like small startups using it all the way to the largest enterprises in the world, rolling it out to their entire customer base to get the benefit of kind of being able to access all of your enterprise data and easily apply agents and capabilities to help you accelerate your jobs. And so we use it across everything from software development to security to driving HR policies. I mean, you guys are notorious for measuring everything about your operation.
Where have you seen the biggest gains? Yeah, it's obviously software development. That's the real answer. The derivatives of that. But yeah, I mean, the speed of software development has been and really product development as a whole, not just coding. So you've seen Actools pick up the velocity of new products. The pace at which we're deploying new products is massively different than it has been in the past. I think you can see this where AWS has always been known for rolling out features really quickly. And we've seen a turbo boost on that in the last year or so as we call them frontier teams, as they think about agentic development as opposed to kind of traditional development.
And it's not code completion. It really is agent first, the agents write all of the code. You're just managing a team of agents and driving that. And it's been fun to see the pace at which innovations for customers has been massive. And it has to be because that's the problem. Our customers out there have an almost insatiable appetite for new capabilities. And that's what we've got to do. Yeah. So organizationally, are you like, do you have any insights on how organizations should change and how? I don't.
I'll say that. Agents manage people, people manage agents. There's going to be lots of people for a long period of time. I do think organizations will change. I don't know the magic answer yet. But we're actively thinking about it. Inside you're just trying various experiments? Yeah, we're trying experiments. We're thinking about pods. As you think about, you know, here's one example is in a product organization, you used to have a team that would own a particular kind of capability for a long time. And you might have 10 people working on that thing.
Well, today you can innovate so rapidly that one doesn't have to be 10 people. It can be three to four people. And they build something so fast, you actually want to move them to different projects and problems. And so thinking about how do you both operate and maintain the things that you built while being agile and flexible to move around on an organization that's as big as AWS is active things that we're experimenting with and playing with. But it's fun and it's enabling for our employees.
They actually love it because they can build faster and do more. But there's a work there. Yeah, it's a fascinating time. And so thank you very much for your time. We could be talking about this for hours together. But thanks for all your insights. Yeah, thank you for having me. And thank you. We love having all your companies as customers and we love learning from them. And I appreciate having me here. Yeah, we'll keep sending them your way. Excellent. All right. Thanks. Thanks for listening to this episode of the A16z podcast.
If you like this episode, be sure to like, comment, subscribe, leave us a rating or review and share it with your friends and family. For more episodes, go to YouTube, Apple Podcasts and Spotify. Follow us on X at A16z and subscribe to our substack at A16z.substack.com. Thanks again for listening and I'll see you in the next episode. As a reminder, the content here is for informational purposes only. Should not be taken as legal business, tax or investment advice or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16z fund.
Please note that A16z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments, please see A16z.com forward slash disclosures. Transcribed by https://otter.ai