1% better

Why Building an AI Agent Is Easier Than Deploying One

Pick one recurring, low-risk workflow where work gets lost across email, spreadsheets, and systems—not just a task inside a single tool. Deploy an AI agent with human review for exceptions, capture every correction as feedback, and expand autonomy only after performance is proven. The goal is not a

59m

Summary published by , updated .

A16Z

Key Takeaway

Pick one recurring, low-risk workflow where work gets lost across email, spreadsheets, and systems—not just a task inside a single tool. Deploy an AI agent with human review for exceptions, capture every correction as feedback, and expand autonomy only after performance is proven. The goal is not a flashy chatbot; it is removing an end-to-end job while preserving expert control where relationship, financial, or operational risk is high.

Episode Overview

Elena Berger talks with A16z partner Sima Amble and LEO cofounder and CEO Vlad Kyle about why deploying enterprise AI agents is harder—and more valuable—than simply building them. Using procurement as the example, they explain how AI-native companies can coordinate work across fragmented systems, earn trust through human-in-the-loop deployment, and gradually take on increasingly complex, judgment-heavy tasks.

Key Insights

Own the Work Beyond the System of Record

Enterprise systems capture outcomes—a price, a delivery date, a signed contract—but much of the real work occurs in emails, spreadsheets, meetings, and specialist judgment. AI-native products gain an advantage when they coordinate that distributed context and complete the full job rather than merely retrieve information from one database.

Exceptions Are Where Automation Becomes Valuable

A workflow that succeeds only on the happy path may create more review work than it removes. Vlad argues that the hard, valuable portion is handling fraud, mismatches, changing requirements, and other exceptions that require context, escalation, and judgment.

Earn Autonomy Through Graduated Trust

Customers will not hand over high-stakes negotiations to a fully autonomous agent on day one. Start with human-in-the-loop reviews, use feedback and evaluations to improve the agent, and expand autonomy first in low-risk or previously unaddressed work before applying it to strategic decisions.

The Last 20% Determines Whether an Agent Reaches Production

A basic internal AI tool can be built quickly, but partial accuracy often still requires humans to check every output. Production-grade deployment depends on integrations, memory, workflows, evaluation systems, exception handling, and sometimes vertical data or specialized model training.

Agents Can Align Both Sides of a Transaction

Buyer and supplier agents may compete on price, but they share incentives across thousands of surrounding tasks: clearer requirements, faster responses, fewer handoffs, and shorter time to market. This creates an opportunity for agent platforms to reduce coordination friction for both parties.

Frameworks or Models

Four Types of Enterprise AI Agents

Sima's model progresses from retrieval agents, which find and synthesize information; to process agents, which execute defined rules and workflows; to policy agents, which apply judgment where definitions are ambiguous; to principal agents, which make broader tradeoffs on behalf of the organization, such as exceeding standard compensation to preserve a customer relationship.

Human-in-the-Loop Path to Autonomous Deployment

Start an agent with expert review rather than full autonomy. Use corrections, feedback, and evaluations to teach the system how the specific enterprise operates; expand from low-risk tasks and smaller negotiations to larger volumes, while retaining expert involvement for complex, strategic, or relationship-sensitive decisions.

End-to-End Autonomous Procurement Flow

Capture demand through a simple input such as a photo, quote, or spreadsheet; check internal inventory and other plants; identify internal or external suppliers; draft and send an RFQ; ingest supplier responses across email, PDFs, and spreadsheets; benchmark and negotiate; then confirm orders, track shipments, and process invoices.

Notable Quotes

"The opportunity for the AI-native startup is to say, we're going to own that entire end-to-end arc."

— Sima Amble

"No company and no enterprise starts with fully autonomous negotiation agents from day one. No one does that. Why? Because they don't trust us and they don't trust the technology from day one."

— Vlad Kyle

"But it's actually only like 20% of the work of the job to be done of the problem. Because 80% of the problem is like, what if the invoice is fraudulent? What if there's like a mismatch? Like what is like if we don't take the happy path?"

— Vlad Kyle

"You can build this in eight hours, and you can build this, but you will only reach 70%, let's say like, the of the performance. And the problem is, 70% of performance or accuracy or however you measure it, it depends really on the task doesn't mean 70% automation, right?"

— Vlad Kyle

"The beauty of AI is a it learns over time. So that's what we're talking about learning loops. And you have the right eval process. You can do more and more complicated jobs over time."

— Sima Amble

Action Items

  • 1
    Map one end-to-end workflow

    Choose a repetitive business process and list every handoff, including emails, spreadsheets, meetings, documents, approvals, and systems. Identify where the system of record shows only the final result while essential context lives elsewhere.

  • 2
    Design for exceptions before the happy path

    For the selected workflow, write down the 5-10 most common failure cases: missing data, mismatches, fraud risk, late delivery, policy ambiguity, or relationship sensitivity. Define what the agent can resolve, what it should escalate, and who owns the decision.

  • 3
    Create a staged autonomy plan

    Start with retrieval or preparation tasks, then add process execution with human approval. Track corrections and evaluation results, and only move to autonomous actions when the downside of an error is low and performance is consistently reliable.

  • 4
    Measure work removed, not demos completed

    Evaluate an AI workflow by whether it eliminates human checking and reduces cycle time, errors, or missed opportunities. If people must still validate every result, focus on integrations, context, and exception handling before expanding deployment.

Full Transcript

Transcript of Why Building an AI Agent Is Easier Than Deploying One from A16Z. Auto-generated from episode audio; may contain minor errors.

If you want to build an aircraft, you need to procure thousands of suppliers. Someone sends a confirmation of like, hey, sorry, this part is going to arrive two weeks later. And if they miss this email, hundreds of millions of them. Procurement historically may have been more in a box. And now it's like, okay, it's touching legal, it's touching finance, it's touching a bunch of different software systems and people. The opportunity for the AI-native startup is to say, we're going to own that entire end-to-end arc. No company and no enterprise starts with fully autonomous negotiation agents from day one.

Why? Because they don't trust us and they don't trust the technology from day one. And by having this human-in-the-loop approach, we are feeding our agent with all the feedback and all the learnings. And then they suddenly trust us for like 10K negotiations, 20K negotiations, 100K negotiations. When you think about what a durable, vertical AI company looks like, what are the qualities that you look for? It's really, really hard to forecast your moat going forward. If you look back at all the best businesses, at the early stages, they were...

If the incumbent already owns the customer, the data, and the system of record, where does an AI-native startup have an advantage? In this episode, Elena Berger sits down with A16z partner Sima Amble and Leo co-founder and CEO Vlad Kyle to answer that question through one of the most complex parts of the enterprise, procurement. They get into why so much of the actual work happens outside the system of record, across emails, spreadsheets, contracts, engineering data, and conversations with suppliers. And Vlad explains how Leo is building agents that can coordinate that context and increasingly take on the job end-to-end.

They also discuss how you earn enough trust to let an agent negotiate on your behalf, why building an internal AI tool can be deceptively easy until you hit the exceptions, and what changes when agents eventually sit on both sides of a transaction. Welcome back to the A16z podcast. I'm Elena Berger, and today I'm joined by Sima Amble, a partner at A16z, and Vlad Kyle, co-founder and CEO of Leo, which builds AI agents for enterprise procurement. Sima, you recently wrote a piece called The Incumbents Are Coming, and it asks a question facing almost every AI application company.

If an established software vendor already has the customer and the data, and the model can work across its tools, where does a startup have an advantage? And we have Vlad here, and Vlad can really help us understand where a startup has an advantage. So, I think a good place to start is the cases for and against the incumbents. So, if a company can connect a capable AI agent to the software it already uses, where does another application actually have a use case and an advantage there?

Yeah, let me back up and frame it up a little bit. So, historically, we thought about there's the incumbent, and then there's the startup, and it's a fight between distribution and innovation, to take my partner Alex Rampel's phrase. However, now there's like this third piece, which we were pointing at, which is the incumbent can layer on one of the models on top, and then there's a much more formidable competitor in the market. So, why do you need an AI-native startup if you've got CloudForce, which is taking Cloud plus Salesforce and putting the two together, and then you've already got all your data and your employees are all used to using the product, so why another product?

I absolutely still think there's obviously still a case for the AI-native startup, and it really centers around the fact that the legacy incumbent is limited to their system of record and that record that they have, and they're not completing the end-to-end job. So, let me put that more concretely in an example. So, say you're a customer, and the customer calls and says they got charged after they got canceled. Resolving that cancellation history isn't just the customer going into the chat and saying, hey, I got overcharged, and the response there, it has to hit billing, it has to look at all the chat history, it has to look at the contract.

That's not one system of record. That's the knowledge around that customer and everything it touched, and that's something that one system of record wouldn't touch. However, the opportunity for the AI-native startup is to say, we're going to own that entire end-to-end arc. So, that could be illegal, so owning everything from brief all the way through trial. Vlad can talk more about procurement, but it's really the concept of owning the end-to-end work. Vlad, do you want to talk about where that does show up in procurement, like where existing incumbent plus a model just is insufficient, and what have you seen just with the companies that you work with?

Sure. So, when we think about procurement, I would assume that most people think about prices, right? So, what is the end price that we negotiated on? And surprisingly, the record looks always very, very simple and very easy. It's just 8K, that's the result, for example. And I mean, it's the same for sales, right? So, even if I come to SEMA and tell them, hey, we now partnered up with another enterprise, look at the signature here for this contract. It looks very easy, but SEMA doesn't see all the work behind it, right?

So, there are probably 30 stakeholder meetings happened, 500 emails, 20 Excel sheets, and that's the same for the counterpart procurement, right? So, you see in your ERP system 8K for aluminum, but you don't see that maybe the supplier did a pushback and asked for 10K. You don't see that a cost engineer had run three weeks of Excel sheets and 3D modeling to find out the prices of the part and everything else. So, that's what we see. So, most of the work in procurement actually happens outside of this ERP or any system of record.

And I'm sure in the workflow that you described, an agent can do a huge number of things and different kinds of agents can do a large number of things, too. And I actually think that's a good bridge into the next question, which is, you know, a year ago, I think a lot of the incumbents were releasing chatbots, and that was kind of the extent of what you would see. But SEMA, in this piece that you wrote, you lay out four different kinds of agents, retrieval agents, which are kind of the chatbots that we're familiar with, process agents, policy agents, and principal agents.

So, can you just walk us through all of them and explain kind of what changes in the kind of judgment that's necessary across all of them? Yeah. So, a year ago, I made this meme, which was the slap on a chatbot strategy, which is essentially like all the incumbents effectively had a chatbot that sat on top of the system record, which you could chat with to retrieve information, maybe do some analytics. And that really was in that first bucket of what the retrieval agent is. And maybe let me walk you through an example of what each of like the retrieval, the process, the policy, and the principal, back to the customer support example, just because it's easier to understand.

So, imagine if you're a customer and there's a service outage, and you're calling in to say, hey, I want to get compensated for this service outage. The retrieval assistant, which the incumbent may have, is going to be able to pull up, yes, this was what the contract term said, and yes, there was an outage, and just verify that information. It's pulling up information about the customer, it's in the database, and it's just sharing it back and maybe synthesizing. The second step in the agent sequence is the process agent.

So, that process agent may be able to pull through an approval of credit and say, okay, based on our policy handbook, it was out from these dates, therefore, we say you are entitled to money, and it can just apply the bill and just do process. There's no judgment involved. Then, as you keep going to the policy agent, so the policy agent isn't just going to apply the process, but it's going to say, okay, in this situation, it was out for 20 minutes, that's enough to be considered a significant outage, and they're applying that judgment because there isn't really a strict definition around it.

And then the last stage, which is the principal agent, you're actually weighing, okay, should we offer more compensation? Because that was a pretty terrible outage, and we want to reserve the relationship, and it's worth doing more beyond even what a process or policy says. But the importance of these four things is that most incumbents started out in, if at all, in bucket one. Now, at least they're marketing that they are moving towards process and policy, meaning they'll be able to apply more judgment. And, you know, if you look at what they've launched, these are workflow agents that will help you get a document signed or input information from a transcription or things like that.

They're very much still limited to, I would say, retrieval and a little bit of process. They have not gotten into more judgment. And I can get into why, like, there are all sorts of incentives that are preventing them from that. But the incumbents are trying, and I think what they're not able to do, they're like, some of them are thinking, OK, let me partner with OpenAI, Anthropic, one of the labs, and try to build out, take that model capability and complement what they have and super power it.

Yeah, well, do you want to say why sort of some of the incumbents are holding back? Yeah. OK, so I would say they're holding back, but they're held back. Yeah, I'm sure they want to be full force, but there's probably two pieces. One, they have this advantage of distribution, right? And they have the customer trust, which enables them to then sell more products with the customer. So take Salesforce. When AgentForce launched, it was very easy for customers to say, yeah, I'm going to sign up for the Salesforce agent, especially if it was offered at almost no extra cost.

And so they have this trust. They have the distribution. It's often pretty seamless to turn on the product. The flip side is there's all these internal incentive issues, right? Which is if you start getting into more complicated agents, it's an internal conflict with an existing product that's offering workflow versus you're resolving the work. And those are two products that are, you're solving a customer support issue end to end versus providing a workflow for a support agent, a human agent. Those are different buyers. How do you, you know, those two teams are in conflict.

And then on top of that, I think from a sales perspective, like what are you selling to the customer? And then oftentimes in a classic like incumbent issue, right, is like there's two VPs, right? And they're different orgs and they're selling different products. And like they will never be able to figure out what the right set of incentives and the right person to sell to. But anyway, so there's all these sort of classic incumbent issues I think that the incumbent also runs into. Makes sense. Vlad, we're across this kind of spectrum of retrieval agent, process agent, policy agent, principal agent.

Where does Leo sit? Yeah, so like we span across like I would say like all of those categories. And it really depends on the complexity and the risk our agents take, right? So sometimes we can like we already have use cases where we run fully autonomously. Sometimes you have the human in the loop. It really depends on like the budget approval, how like how much, how complex it is, how risky it is. But maybe like coming back to what Sima said about trust for the incumbents, like you talked about like internal trust.

I think there's like also an external trust thing, right? So like how do you convince someone to go from, hey, just like an agent that retrieves some information and maybe runs processes to do like something completely autonomously? It's like obviously there's like a product component to it, but mainly it's a people component. So they need to trust you. And like we as a startup scale up, we have to earn the trust. Incumbents already have to trust, but this also means they like they can destroy the trust if they ship something like too early and the product maybe doesn't work or it's bad or it's like works decisions in a bad way.

So that's what we can do. And interestingly, what was like really surprising for us, like those process agents actually became kind of like a side quest for us. Because it's like when we talk about when we look at the invoice process, like invoice agents as a process. So you retrieve some information, you match this across like other documents, and then you push this back to SAP Oracle, like a very clear process. And then we figured out, okay, that's like 100% of the software market, right? So that's invoice software.

That's how you build it today. But it's actually only like 20% of the work of the job to be done of the problem. Because 80% of the problem is like, what if the invoice is fraudulent? What if there's like a mismatch? Like what is like if we don't take the happy path? And so we like very fast shifted to the next step of agents like doing those exception handlings. And we convinced the customers by, I think we got like lucky being like always slightly ahead of the curve.

So we were able to pitch the next generation of agents. So for example, like when we started out three years ago, it was like just retrieving a document. Like not impressive at all today, but like three years ago, this was like crazy impressive. So we pitched this to customers, we find out, okay, that's a real problem, that's a real use case they would pay for. And then we were able to ship this some weeks later. And in the same way, we are doing this now for the next step for process agents and for like fully autonomous agents and for long running agents.

We are pitching this to them, finding out it's a problem, and then we're able to ship this very fast. I think an interesting point on trust is this internal versus external trust. The other lens on that is, yeah, so you need your customer to buy the procurement software and trust to use it for their internal processes. One of the really interesting things when we first met Vlad was that they're also doing the negotiation. So you have to trust that the Leo agent is going to then interface with a third party.

And there is that piece of trust. And of course, I think a lot of people feel burned by the incumbents and like, you know, they're pretty limited and haven't been able to do what they've marketed they've had in the past. But putting that aside, like, I don't know, maybe Vlad, I'd love to hear a little bit how you convinced the customers to trust an AI agent to now take on negotiations. Yeah. And when you do that, can you also just like paint the picture of like what's involved in procurement and who are your customers and what kind of sort of legacy systems are they used to?

Yeah. So like when you think about procurement, you like maybe just think about like purchasing or like when you look at like B2C world, like just you buy something that actually is like a very, intense process which runs the economy. And it includes multiple stakeholders and a lot of departments. Legal, cost engineering, procurement, finance, all of them have to work on these decisions. I think the reason is how we convince them is like what I already said earlier on. So we were always a little bit ahead of the curve.

So we knew the technology is coming. Even before Chattopadhyay came out, we just started a few weeks before this Chattopadhyay breakthrough. So we already had all of the problems. Then we had the hype in the tech bubble, and we were able to pitch to enterprises and we quickly figured out, okay, this could be an interesting use case, like just chatbot applications or retrieval agents, document processing. And then we were able to find out the problem and then ship this quickly to them. And then obviously, this is the people factor of trusting.

So we are telling them something about and we're able to really ship something in production. But then there's also this product perspective where we have a lot of evals in place. So no company and no enterprise starts with fully autonomous negotiation agents from day one. No one does that. Why? Because they don't trust us and they don't trust the technology from day one. So we have a very easy approach of having a human in the loop. And this is extremely helpful for us because we have the perfect negotiation agent, which is overall the perfect procurement negotiator.

But we don't know exactly how a Fortune 10 enterprise, like the specific Fortune 10 enterprise operates. And by having this human in the loop approach, we are feeding our agent with all the feedback and all the learnings. And then they suddenly trust us for like 10K negotiations, 20K negotiations, 100K negotiations. And then you also have like other agents that are like more long running. So when we talk about like multimillion dollar negotiations where you analyze complex 3D models and technical drawings, there we on purpose have always experts in the loop.

So there's an agent running for multiple hours and then we ask for feedback of the cost engineer and then it does the next work and so on. Maybe just to double click on that, where do you put the human in the loop on the negotiation side? You mentioned the cost engineer, but like if you were going back and forth on a deal, is it mostly around the data for something like cost engineering or is there anything else where you have humans in the loop there? So again, it depends on the level of negotiation.

So we have to distinguish between negotiations where you just like a negotiation where enterprises, they never did those because they didn't have the capacity. But by deploying agents, they can just capture savings that they weren't aware of. So before agents or before Leo, they just didn't care about everything which happened below 50K. So you can just like this is maybe a hack for like other startups, you can just send an enterprise an invoice for 40K, they will probably not negotiate because they don't have the capacity to do so.

Except they have Leo agents, then we are going to negotiate against you. But other than that, they are just like just paying that. And obviously the risk of like, you didn't negotiate at all. So what is the risk now of having a bad negotiation agent? Nearly zero. So maybe we miss out on some negotiations, but it's better than nothing. But still, we're talking about business relationships and business relationships are not always about the cost and the money. So maybe you're not spending a lot on a vendor, maybe you need this business relationship.

So a good example might be like podcasts or marketing services. That's probably like a friction of the spend, but you have a clear business relationship with someone setting up the studio and you don't want some random people doing that because they already know how A6nZ operates and how you want to record all this stuff. So there we have human-in-the-loop approaches, where procurement people care about the relationship. So they care about the voice of tone and how it works, but mostly autonomously. And then we have the other set of agents where we always have a human-in-the-loop approach.

And it's like there are multiple steps in negotiating. It's not only the price. It's also like, how is the contract design? So we're talking about legal. How is collection design? So we talk about finance, obviously, like cost structure. So we're talking about really like cost engineering. Then we talk about commercials. That's procurement. And those are not back office people. Those are like highly trained people where they have very specific knowledge of a very specific process of a very specific company and very specific industry. And they feed those long-running agents of Leo with those insights.

That's another example of how procurement historically may have been more in a box. And now it's like, okay, it's touching legal, it's touching finance, and it's touching a bunch of different software systems and people and both specialists and more generalists. Yeah. Can we map this onto a specific customer? Like not a specific customer, but a specific vertical? Like I'm a drone manufacturer. I manufacture humanoid robotics or something like that. Like how many parts do I have to order and procure? How many factories am I touching?

How many suppliers am I touching? Just all of those things. If you want to pick, maybe Vlad, a vertical that is just managing all of this complexity with Leo and just kind of take us through what their experience is, I think that would really help just illustrate exactly just everything that you touch. Again, when we as Leo, when we talk about procurement of purchases, we don't talk about laptops and pencils. We think that's solved also by Leo agents, but that's easy. We solved this like three years ago.

We talk about when you want to build an aircraft or robots or drones. We're now doing a podcast about AI hype, but even AI needs to be built. So you need data centers. And building means procuring. Someone needs to, if you build an aircraft, you need to procure thousands of suppliers. You need to build a factory to build this airplane. And really small frictions can have a crazy impact. If you're running a very large project of building a data center or building an aircraft, if there's one specific part which arrives two weeks later, this can have a damage of hundreds of millions of dollars and postpone the project.

That's why all of those decisions have to be coordinated. And one part is you need to figure out what you need, with what suppliers you work, what is the best supplier to get this part. But then once you decided all of this stuff, there's all this operational back office stuff behind it, which might sound unnecessary and boring. But again, operational means someone sends a confirmation of, hey, sorry, this part is going to arrive two weeks later. And this is one of 500 emails in the Outlook or Gmail of a procurement manager.

And if they miss this email, hundreds of millions of damage. And this happens regularly, because the only thing they store in their system record is then just the date. Not this Wednesday, next Wednesday. That's what you see in the system. But you don't see, maybe that's okay, but maybe that's a $100 million damage. And someone has to decide that. And that's also what our agents are doing. They are not only retrieving the information, they are then making the decisions. Does this have an impact? What kind of impact?

And how can we resolve this? Yeah. When you are sort of so deeply embedded in the physical world, what kinds of physical world problems can you intervene with? Some things I would think are just unsolvable. Let's say you have a shipment coming in and a bunch of stuff falls off the ship, or the street is closed or whatever. There are things that you can't do. And obviously, there are things that you can do. So where can you intervene? And where does that really make a difference?

Actually, it's about probability. Obviously, you can't change if there's damage on the ship. Every example that you manage, you can't change that. But if you have all of the context, you can predict that. Because you can predict how reliable is the supplier. So there are ways to protect the goods that you're shipping. And if you have all the context, you have one supplier where 20% of the goods are missing, and then 1% of the good is missing. And maybe this one with 20% is 10x cheaper.

But for this use case, it's fine for you to pay 10x the amount because you have a higher probability that this thing actually arrives. And this is the powerful thing because you have context not only within one enterprise. So we talked about multiple stakeholders, but there's also context on the outside world. So the agent should have context of all the news out there. Maybe even having context of some bad sale on Polymarket. Okay, those disruptions are going to happen. Then information about on the supplier side, on the seller side, on the demand side.

And by combining all of those contexts, I wouldn't say that there is a limitation in the long run. Obviously, today, we have different sets of probability, but we can help throughout the process. And this is what we are building at Leo. So it's much bigger than just procurement into a company. It's more like intra-company and how enterprises are doing business with each other, so buyer and supplier side. You described Leo as a multi-agent system. So can you describe what the different agents are doing? One level is that Leo agents span across all those four categories that Sima mentioned in her article.

And it depends on the risk and the complexity. So we use all of them, so multiple agents. But the other thing is that in order to do a job end-to-end, those agents need to share information with each other. They need to do this in a very specific order. And when we talk about a multi-agent system, this is essentially what we are doing. We are solving the task end-to-end. And because the human level task involves eight people, eight stakeholders, and maybe three departments and five different software tools, we need to cover all of those to do the job end-to-end.

And those agents need to then communicate with each other. And only with a multi-agent system, you can do a job end-to-end. When we started out, we started off with a retrieval, more like a co-pilot, obviously, three years ago. But then the next step was a single agent. But then we very quickly discovered, okay, you can't solve in negotiation even without having a contract agent, without maybe having an agent looking at the news and everything that I described before. So that's what we define as a multi-agent system.

So if you have a bolt, like an airline company needs to procure a bolt, Boeing needs to procure a bolt, what exactly is that process for procuring the bolt? And where does Leo step in on that process? Yeah. So this is one of the purchases that we can run fully autonomously. And we can do this because of this multi-agent system. So first of all, someone has a demand, right? So they need to somehow communicate it. And even this part is extremely complicated. So you need to call someone.

Maybe you open up your laptop because you're a construction worker, you open up your laptop only every second week, and now you are required to work with SAP or any other ERP system. So you can't even issue the demand. So this is how we make it very easy. So you take a photo, you upload a quote in Excel sheet, and it's actually everything that you should know about procurement. No one cares outside of the procurement department about categories, GL accounts, framework contracts, no one cares. We in the procurement world care, but no one outside that cares.

And then our agents take off and they check the inventory. They find out, okay, they ask another plant, okay, can we source those bolts internally? No. Okay, then I'm calling the, I'm talking to the sourcing agent, finding out, do we have internal suppliers? Do we have external suppliers? Then another agent has to draft the RFQ, send out the RFQ over email. Then a bunch of emails arrive. Some of them are completely nonsense. Some of them are just in the email, some of them are PDF, some of them are Excel sheets.

We retrieve those informations. Then we do the next step. Oh, maybe based on our price benchmarking, there's an opportunity to negotiate. And then we have agents that essentially decide on the next step. So negotiation could mean strategic negotiation with a human loop. This could mean autonomous negotiation. This could mean auctions and e-auctions, calling them the specific agent, doing the negotiation, and then doing the It's like end-to-end, think about confirming the order, shipment tracking, invoices. And we are able to run this fully autonomously, capture all the context.

And then obviously the next powerful thing is do this for more complex parts where we talk about direct procurement, where we also operate. What's direct procurement? So everything I just described, the main goal is automation, so you can run this process fully autonomously. And then throughout the process, you can generate even more savings. So we look at it like, okay, what is this end-to-end job to be done? How does it look like? So what are they doing a thousand times a day, but actually they want to do it zero times a day?

We've run fully autonomous agents. But there's also opportunities of what are they doing zero times a day, but if a business would do this a thousand times a day, that would have a crazy P&L impact. Autonomous negotiations on spend they never negotiated before. And this is in the indirect procurement part, think about MRO parts, building a factory, the bold example that we did. But also laptops and pencils, marketing services, someone who needs to build up this podcast studio. Those are all indirect. And then we have direct parts, just like when you build an airplane, those are all the suppliers that you actually need to build the airplane or to build the drone or to build the robot.

And then we don't talk about 50,000 suppliers, we talk about 100 suppliers or 2,000 suppliers maximum. And those are extremely strategically important. And you have maybe on one supplier, one billion of spend. So you don't want to run an autonomous negotiation. You want to run a negotiation which takes three months and where you're crazy prepared and where you have engineers on your team analyzing, okay, what's the indice for aluminum? What's the indice for oil? How did the price change? So you really take over all of those drawings, you check the quality of this part.

And this is where it gets really exciting, deploying agents. Yeah, and for something like that, presumably, you'd have the expert engineers and the other procurement people kind of more as the front of house and the agent is more back of house. Is that the idea? Or is the agent like, actually, it's like you sit across the table and you're shaking hands, and it's like the robot instead of the human who's negotiating? So yeah, is it more back of house or is it still front of house?

It's obviously more back of house because you need these complex multi-million dollar negotiations. And that's, again, a beautiful example. 90% of the work is preparation. The end result that you see in your system of record is like, oh, instead of like 1 billion, I paid 900 million dollars. Okay, but there's like three months of preparation and like 10 people working full time on that. And obviously, this is happening in the back. But actually, we have some use cases where it's also helping in real time. So think about, let's assume we would now have a negotiation and I have a perfect preparation.

Same as like with those notes that we are having here. Imagine like while we are negotiating, I would have like real time insights on my screen popping up where you tell me the index for oil changed 10%. So like it's increased by 10%. So that's why we need to increase the prices by 10%. And I would have like an initial, like an immediate pop up of like, that's true, like oil increased by 10%. But the product has only 30% of oil contained. So like you shouldn't increase the price by 10%, but maybe only by 4%.

So yeah, they are like exciting use cases also like in the real life. Yeah, we know that, you know, companies like Harvey and Decagon are really fine tuning models now. What kind of underlying models do you use and how do you approach things like fine tuning or harnessing? Yeah, so like we believe like you can, so like we use multiple models from like all providers. And we really see this as a like, obviously like as a commodity, right? So they like have really good like general business purpose or like reading, creating a PDF and like creating an Excel sheet, like all of this stuff.

But we also believe that like, for some use cases, you can get extremely far with like combining the foundation model with a harness. And you can maybe reach like 100% of like the job to be done. But there are also some use cases where you can have like the best foundation model, the best harness, whatever that means, but like the best harness, but you still can get only to 80%. And like, it would say like, good examples for that is like, for example, what, when we talk about negotiations, what like cost engineers doing, right?

So they're like analyzing drawings, and then they like defining, okay, what should this, this part actually cost like, and that's why it's called like short cost modeling. And there's, there's definitely like an opportunity where we like thinking and already started fine tuning the model to get them to, to get them to 200%. And this part, another example is like price benchmarking, where, think about like the, like, you would have a quote, and in a perfect world, you would just drag, drag and drop the quote somewhere, and you would get the perfect price.

But it's like, and all those, like all those informations, they're like not publicly available, right? So there's all like, those are like all preparatory data based on like one, one enterprise, like across multiple enterprises. So like general purpose models can't train their models on that. So what we are like, what we are thinking about is like, maybe like not training just an LLM. But I think like, what we see now with models like Jeff, or so popping up where you have like an, you train it on like, text data, but the out like the the output is actually like an outcome, or just like the perfect price.

And you can't do this with Harness, because, like, if you would give me like an, a quote from BCG and a quote from McKinsey, they, they could do like the exact same work, but this could be like a 10x different price. And I would have like no idea, like, what is better. But if you give this to a procurement manager, he would like initially have a gut feeling. Or like, okay, like this quote makes sense. Um, like, I think, I think, like good examples, again, like content creation, like, always like, I don't know, like how much I should pay someone for creating a video.

But it's like a gut feeling behind it. If I ask like another video creator of how to, how to do that. But if you ask them to write down the rules, they can't do this, because it's like, just like gut feeling and instinct. So, and that's where we see a lot of opportunity, actually, training an agent, but not maybe like a classic LLM, but more exact, like, again, like what, what companies like, like models like, like JAV are now doing on the outcome based. So we have like price benchmarking, short cost modeling.

Sima, we've talked a little bit about how labs are really moving into industry specific work or working with incumbents to do this. Um, when you think about what a durable vertical AI company looks like, what are the qualities that you look for? One is around, you know, owning the end to end work that we're talking about building up this data asset and being able to do something that hasn't been done before in many cases. This is all said, I think when you talk a lot about moats, it's really, really hard to forecast your moat going forward.

If you look back at all the best businesses, at the at the early stages, they were, they were just thinking about, okay, I'm winning customer trust, I'm selling more to them. And there's a lot of opportunity versus, okay, I'm going to do these six steps and then get to the seventh step, and then we'll have a moat. And so I think we we talk a lot about defensibility and durability. And I think part of that is you're locking in the customer, they, there's more dependencies, they find it valuable, and you're doing more of the work.

And here, it's, it's truly like, okay, you know, old CRM company was a log for all the deals, new sales, AI agent is actually owning a lot of the sales prep process and the outbound process fielding inbound and doing a bunch of work. So the company overall customer is dependent on that product. And that's like a really important signal of getting to the moat and everything we talk about in terms of stickiness and network effects and all that is sort of downstream of, of that initial like customer use and the value of the product.

Vlad, have you had conversations with customers or potential customers who ask you, you know, why should I buy your product? Why can't I just, you know, plug into a model and do this myself or like use whatever existing system of record I have plus a model? Like, like, what what do you tell them? And how do you convince them to use Leo? Yeah, 100%. And that's a very fair question. Right. And like, even if you look like internally at Leo, so like the, the first use case three years ago, which like kind of like went viral in the procurement world was like, just like having a quote and then getting this information into SAP.

So like very like, again, like, technically, like, but like, like tremendous business value. And so you, you have like this retrieval agent, getting like all of the information, putting it into SAP. We like this was our like, first product. And we had an engineering team, like obviously small, just like the three of us, or maybe like four people building it, and then selling it. But this is nowadays, a case study. If you if you're applying to work at Leo, so we give this to people to like to build this and they have like, eight hours to do so.

So what I want to say by that is like a product that we like one of our first use cases can now be somehow built by engineers within eight hours. So because it's very easy to build stuff nowadays. So obviously, there's a question, okay, well, so someone can build this within eight hours. Okay, cool. But then couldn't like just procurement departments also just build everything in two months? And the answer is like, yes, you can build this in eight hours, and you can build this, but you will only reach 70%, let's say like, the of the performance.

And the problem is, 70% of performance or accuracy or however you measure it, it depends really on the task doesn't mean 70% automation, right? So it's this can be mean that you're like, have 70% of the performance, but you still need to do 100% of the work. Because 70% is not not that much. So again, like all of the people have to have to check the data. So maybe you even created like more work, more work than they did before. Two things to layer on to what Vlad just said.

One, overall, it's good if there's more adoption of the base models or just like GPT products, because it means that people are also willing to trust vertical specific products as well. So I think that increased familiarity, comfort, excitement about AI tools is generally just good for the market. The second thing is, I was chatting with the management team of a Fortune 500 company two or three weeks ago. And one of the things they mentioned was they had tried to build out their own cash collection product was a big enterprise business.

And they like after, I don't know, three or four months of work at a minimum, they had found that there wasn't enough context, there was poor quality context. They had a lot of recordings, a lot of screen grabs. They're trying to pull it all into one system. But they there was no there wasn't good enough. And then there was this giant question around, OK, like, you know, we've got now two different ERPs and we are about to acquire another one. Who's going to update all the mappings, test it, you know, test out?

OK, does it work? And then like we keep talking about exception handling. You now need to map that onto a totally different system, a different way of doing things. And I think they quickly realized that the internal build didn't make sense. And so we keep hearing stories of this where people are like, OK, I'm going to do the internal build. And they're like, wait a second. It's not different from what DIY has ever been in the past, which corporates have always tried. But I think enterprise companies generally realize that there's their core competency and then there's building internal tools and they should focus on the first camp.

Yeah, yeah, makes sense. So that's exactly what I what I've meant with like, obviously, they can get to 80 percent, but those last 20 percent really matter. And you can only get like they matter to get into production. So that's why you need like this harness, right? So you need all those like integrations, memory, workflows, and sometimes you need vertical data to do that. And like the Pareto principle, so like those 20 percent can make like 80 percent of the of the effort, like they are making 80 percent of the effort.

So like to all like those Fortune 500, 150 companies, so you can do this. But then let's say like procurement workforce or AI procurement agents should then become like one of your core competencies. And you should evaluate whether this makes sense or not for you to have this in your core skill. Vlad, I'm curious if you're seeing suppliers start to use agents or AI at all and kind of like what happens when both the buyers and the suppliers are fully AI enabled. Yeah, so like we 100 percent believe that like in the future there will be like agents on both sides.

And this this makes so much sense. But like surprisingly. Certainly what we see is like, when we look at the supplier side, this also kind of like equals the seller side, right? So what we see is like that the sales side was like always ahead of the procurement side. But what we are now seeing with suppliers for this Fortune 500 companies, this actually is not true. So they are like maybe advanced in like, let's say, video recordings and like using tools like Granola and all the stuff, but not like really having agents deploy, they're automating the work.

And the cool thing is, procurement is unsexy, right? So sales are sexy, procurement is unsexy, but it's like one process and procurement is the counterpart. But now the good thing is, in those industrial companies, Fortune 500 companies, procurement has the bigger power to the supplier. Because you as a typical use case, like automotive supplier, you dictate to your suppliers what they should use, what the quality has to be, how they have to answer to a specific RFQ. So the opportunity is now, if we serve the procurement department and they can dictate what a supplier should use, why aren't we also pushing them to like Leo agents that are also helping them automate the work, owning then both sides of the transaction?

And Elena, I know we were talking earlier today about how can you have two parties on the same platform? How does that work? I think it could even extend into like legal work, right? And these are two very adversarial parties, right? But if you have two law firms with clients with different interests, but both benefit from knowing, okay, here's the latest draft, here's where we are with the open issues, here are things that have been agreed upon, and just even tracking that, that doesn't really exist right now, right?

That's all being created by humans. And so that coordination effort, an agent could be doing. Well, it's super cool to think about how both sides are kind of maybe evolving in tandem. One side might be going a little bit faster, as you're talking about, Vlad. But over time, potentially people are just on the same platform and actually it's way better coordinated for everyone. 100%. Because also like what we mentioned in the beginning, right? So the obvious question is, okay, but we also talked a lot about negotiation agents.

What if both parties have negotiation agents? And price is 100% the point where it's like a zero-sum thing. They have different interests. But what we also discussed is that price is the outcome of 5,000 other tasks that happened. And on those 5,000 other tasks, they have the same incentive. Sales wants to have as little friction as possible. Buyers want to have a really fast time to market, right? Again, like coming back to building aircrafts, building data centers, you want this data center to be built as fast as possible.

You don't want to be building like six months later just because it takes so much time to analyze all of the responses from suppliers. And you want to make the sale also fast. So like all of those 5,000 other tasks, the incentive is exactly the same. And that's why you can deploy or Leo can deploy agents also on both sides, automating the work of all those other tasks. This is the beautiful thing. I think that there used to be this logic of like, you shouldn't customize your software too much to one end user, one end customer.

But I think something about LLMs and AI in general is like, it might increasingly be possible to customize without slowing yourself down as a business too much. So I'm curious if that is something that you're seeing, Vlad or Sima, and kind of what does that mean for end buyers of software? I think the overall principle, there's a lot of forward deployed work happening right now. And part of that is because the state of the customer data and understanding customer N is a lot harder than understanding N plus one.

And so we're sort of in the early stages of deployment overall. And that's why there's still a lot of humans as part of this product. And by the way, that is something that is harder for the incumbents to do because they're also like not set up in a way to have even the way their product feedback cycle works where they have a they have implementation teams. That's very much an afterthought versus something that feeds into the product side. The beauty of AI is a it learns over time.

So that's what we're talking about learning loops. And you have the right eval process. You can do more and more complicated jobs over time. And part of that is automating the deployment itself. And so you can be curious to hear how Vlad is doing it. A lot of our companies even are doing that at a rapid clip where more of the customization A is being handled in an automated way and B, the customer is able to turn the knobs and levers around customization via software versus, OK, I needed to bring in A, the original it was like, you know, you brought in Accenture to do your SAP customization.

And now it's like, OK, I've got a forward deployed team that's going to build some help and spec it out. And ultimately, it's going to be completely software driven. Yeah, I mean, that's the reason why when you look at the org structure of Leo, like 85 percent of the people are engineers. And even if you look at the people where they don't have an engineering title, they have mostly an engineering background. And the reason for that is we obviously don't want to be a consulting company, right?

So we make sure that we have overall the best agents in indirect and direct and finance in those parts. But then, as you mentioned, there is a lot of forward deployed work to do if you go to enterprises because they have different nuances in their processes. But how we work is, as you mentioned, building a product in a way where we're reducing this customization, but also where it's a lot of self-service. So the job of our FDEs and forward deployed engineers is, on the one hand, making it self-service, but internally it's like automating their own job, right?

So their KPIs, you are seeing this happening multiple times. Literally, your job is to automate yourself. And then if you automate yourself, you go to the next task. I think it's also an approach that Google is doing, but I 100 percent agree with you. And that's exactly what we are building and how we're doing it. There was recently a very big system of record event. And we're not talking about Dreamforce. We're talking about the Bots and Buyers Summit that Leo hosted in New York. And just wanted to kind of hear stories from the ground and what you're seeing among buyers.

What are people excited about? What are people looking ahead toward? What are people asking you for? Can you just kind of tell us some stories from that day and that event? So this was the third time that we are doing this event. Now we did it in New York, just a few blocks from our office here, and over 100 procurement leaders arrived. And we made sure that when we do such events that we only invite high caliber people. C-level, CPO, vice president. And there are two things very different of how we do this or why they are so amazed.

So the first thing is, when we look at how procurement used to work in the last 26, 27 years, a lot of tools emerged. So there are all those technology landscapes and you can find them on LinkedIn. You will see there are like 500 procurement tools. And this is also the reason, but when you talk to procurement people, it's very painful. And I challenge someone to find someone who loves to work with procurement. No one does. You can really state people hate working with procurement. And I'm talking about the requesters and I'm talking about the suppliers.

And even people in procurement hate working in procurement. So what's going on if there are like 1000 tools? And the reason for that is like all of the tools. They just made the process more efficient. That's all. But they never changed how those people actually work. And it's crazy to see that they work in emails and on Microsoft Teams and in Excel sheets and PowerPoints. It's like their main channel where they work on. And it's like zero AI enabled. Obviously, in this part. And the other thing is like we give them like a very cross department perspective on procurement.

So we are not talking about like, look at this crazy invoice feature that we developed. But we're more looking like someone needs something. And in the end, you have it on your table. And this can be like across indirect, direct, logistics, finance. And you can see how we are doing it here. And we also putting it into like more into like a physical world, because like AI agents, it's very abstract. So what we are doing is we building a booth. And we even have this in our offices.

And also in New York, where you can walk through the booths and experience like all of those agents like really hands on. And that's what the people love. I mean, the next event will be with around 700 people. And so you can imagine how crazy this grows. Procurement people gone wild. Yeah, yeah. It's gonna be that one is in Munich. Yeah. So we're doing them like in Europe, Munich and in New York all the time. If you're listening and in procurement, you know where to go.

Maybe, I guess one question, one question for me. What do you think it takes to get people excited about procurement? Is it the agents? Is it the people? Is it the time saved? Like what or like something else like you mentioned it like procurement is one of these things that I remember, you know, people aren't, they don't like it. Like it's like a universally kind of disliked low NPS area. I remember talking to a guy who was out of procurement, like, I don't know, seven or eight years ago as I was looking at this category.

And he was like, oh, I hate talking about this product. It's like I use, you know, this legacy system of record. I'm on Coupa and I don't, I don't want to buy anything else. I don't want to talk about it. It's fine. It was like the most disgruntled customer call I've ever done out of like millions of them. But I mean, I'm curious. Yeah. What it is that you think, you know, really gets people excited about this category? Yeah. So I mean, it's like, it's like, and this is also why I like procurement.

It's like, it's on one hand, like, so like the reason like why we started in procurement it's not essentially like what happened, but like how people react to it. Right. So if you talk to the people, they're like really frustrated. So this means it's like highly emotional topic, but it's like, like, let's be honest, like B2B SaaS. Okay. But it's like highly emotional. So that's a good thing. And then when, if you combine this with like something which is boring and niche, this is also an advantage because again, like the, it's, it's like also easy or easy for us to amaze those people.

Right. Because like the really last revolution they have seen is like 20 years ago. And then maybe a nicer user interface 10 years ago, but nothing else happened. And so you have like boring, highly emotional, and then plus crazy business impact. Right. So like, it feels like it's unnecessary, but like, I told you like some examples. So like, it has like obviously crazy P&L impact, but it has impact on like the whole economy. Right. So we're like, we are talking about like how data centers are built, how aircrafts are built, how cars are built, how drones are built.

So it's extremely important that you have a fixed procurement process, not only to like make it happen and build something. But then also when you talk about like when we, when you look at like the competitive landscape. So to get like 1% margin increase, you need to make 10% more revenue, 10% more sales. So like if you just managed to get like 1% savings, it's like equals like 10% of sales that you have to do to get the same outcome in your P&L. So it's tremendously important.

And like you combine all of those three things and then you have like a trillion dollar business opportunity. That's, that's my opinion. Like for procurement, but there are probably also like other things that like emotional, boring, and have a crazy business impact. Well, Vlad, thank you so much for joining us. This was a ton of fun. Thanks for listening to this episode of the A16Z podcast. If you like this episode, be sure to like, comment, subscribe, leave us a rating or review and share it with your friends and family.

For more episodes, go to YouTube, Apple Podcasts and Spotify. Follow us on X and A16Z and subscribe to our sub stack at a16z.substack.com. Thanks again for listening and I'll see you in the next episode. As a reminder, the content here is for informational purposes only. Should not be taken as legal business, tax or investment advice, or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast.

For more details, including a link to our investments, please see a16z.com forward slash disclosures.

Go deeper