What Would Make an AI Assistant Worth Paying For?
Audit one recurring administrative task this week—such as price-drop refunds, inbox triage, HSA reimbursements, or calendar coordination—and delegate the lowest-risk portion to an AI assistant. Start with work that saves time or money without changing your commitments. Require the assistant to draft
50mKey Takeaway
Audit one recurring administrative task this week—such as price-drop refunds, inbox triage, HSA reimbursements, or calendar coordination—and delegate the lowest-risk portion to an AI assistant. Start with work that saves time or money without changing your commitments. Require the assistant to draft or notify you before any irreversible action, then expand permissions only after it earns trust through consistent, accurate results.
Episode Overview
Anish Acharya and Assistant Bench creator David Pollan explore what would make personal AI agents valuable enough for everyday consumers to pay for. They argue that winning assistants will be proactive, nearly invisible, and focused on eliminating administrative burdens rather than merely making users marginally more productive. The conversation also examines voice interfaces, specialized agents, agent-to-agent commerce, trust boundaries, and business models.
Key Insights
The best assistant is the one you barely manage
Consumers are unlikely to adopt agents simply to become 10% more efficient. The compelling experience is an invisible assistant that completes tedious work, reports the outcome, and does not create another workflow to supervise.
Prioritize cost-saving automation over generic productivity
High-value early use cases include recovering flight-price credits, filing HSA reimbursements, and adjusting home sprinklers based on weather. These tasks feel like found money or reduced hassle, making their value clearer than abstract productivity gains.
Proactivity is powerful—but trust has a hard boundary
An agent can safely take proactive action when it creates upside without changing the user's commitments, such as drafting an email or obtaining a travel credit. Actions with meaningful consequences, such as switching insurance, should require explicit approval; one major overstep can destroy user trust.
Voice wins when hands and attention are occupied
Voice-based agents are especially useful during activities like biking, cooking, gardening, or commuting, when a screen is inconvenient. David describes using ChatGPT Voice to organize email, send replies, and create calendar invites during a bike ride.
Agent-to-agent commerce will reshape digital marketplaces
When agents book restaurants, compare products, or transact for users, marketplaces built around ads, impulse purchases, and human friction may need new models. The speakers expect new infrastructure for agent identities, security, auctions, recommendations, and supply-demand matching.
Notable Quotes
"You hire an assistant to allow it to be invisible, right? The best employee I've ever had are the ones that are just doing their work, getting things done. They update you. Here's what I did. And you're like, wow, amazing. You don't hire someone because you enjoy managing them."
"I think there is massive defensibility around proactivity. I think that's actually going to be something that separates the winner from the, you know, the meat of the pack. The agent that can be most proactive, that can serve as an invisible assistant and start doing things on your behalf without you even asking."
"It's not about cost reduction. It's about possibility expansion."
"It opens the door to a truly consumer aligned Internet, which I think we have been lacking for quite some time."
Action Items
-
1
Choose one invisible-admin workflow
List the recurring tasks you dislike but must complete, then select one with a clear outcome: organize incoming email, identify lower prices after a purchase, prepare forms, or coordinate scheduling.
-
2
Set an autonomy threshold
Separate tasks into “can execute,” “can draft,” and “must ask.” Allow automatic completion only for reversible, low-risk actions; require confirmation for purchases, account changes, or commitments made to others.
-
3
Try a hands-free work session
Use a voice-enabled assistant during a commute, walk, household chore, or workout. Ask it to triage messages, capture tasks, draft follow-ups, or review your schedule, then verify the outcomes afterward.
-
4
Measure value in money and annoyance removed
For one week, track dollars recovered, minutes saved, and tasks avoided—not just output volume. Keep using automations that create clear economic value or remove disproportionate friction.
Full Transcript
Transcript of What Would Make an AI Assistant Worth Paying For? from A16Z. Auto-generated from episode audio; may contain minor errors.
The general population does not care about being 10% more efficient. I have a hot take thesis that this muse charm is actually less about trying to win the hardware game and it's more about data collection in the real world to fuel Zuckerberg's future metaverse of mapping out the actual world. We were joking internally like we're days away from an agent messaging someone saying I noticed you weren't that into her so I went ahead and broke up with her. There is massive defensibility around proactivity. I was biking to work.
I want to get stuff done and I just start talking to chat GBT voice and throughout my 30 minute bike ride to work, categorize my emails, submitted them to different labels, sent out calendar invites, got to the desk, inbox zero. But it will be a very fine line because if you cross that line once, you lose all trust with your user. It does feel like there's going to be infrastructure and products that exist only for agent to agent interactions that doesn't exist today. What do you think?
Is it profitable that we're going to see personal AI agents have gone from experiment to one of the fastest moving areas in consumer AI? But what would it actually take for one to become part of everyday life? In this episode, A16Z's Anish Acharya sits down with assistant benchmark creator David Pollan, who has been testing dozens of personal agents across everything from email and travel to shopping and financial admin. They discussed why the best agent might be the one you barely notice, proactively handling the tedious parts of life without needing to be managed.
They also get into how much autonomy will actually give these systems, whether we'll interact with them through text, voice, apps, or wearables, and where specialized agents could still beat the big general purpose players. And looking further ahead, they explore what happens when agents start transacting with other agents, reshaping everything from commerce and restaurant reservations to how we spend our time online. All right, welcome to the A16Z show. I'm so psyched to be here with my friend, David Pollan. We hung a couple of weeks ago. We're both so enthusiastic about everything that's been happening with personal assistant and consumer AI generally.
And man, so much has happened in the last couple of weeks. Maybe to tee it up, I sort of think that there's been a few moments where some part of the ecosystem has sort of seen God. And ChatGPT was a big one, I think, for a lot of folks in November 22 and a bit of 23. It was like, what is this thing that can write emails and poems and start to have a conversation back and forth and feel like you're interacting with a synthetic person?
That was, of course, the beginning, the big bang. The second moment really was coding agents and everything that developers have been obsessed with around sort of quad code and codecs and the ability to trivially create software. And then finally, most recently, with Instinct, Muse, and ChatGPT work, we've had this extraordinary enthusiasm around personal agents and what they can do for us as consumers. So we're here to talk about all things personal agents. We're going to go as deep as we can. We're going to assume that everybody has seen many of the same things that we have seen, and we're going to make this the first in a series of conversations.
Welcome, David. Thank you. Thanks for having me. Excited to be here. Let's like zoom out and maybe talk about the arc of personal agents over the last few months, the last few weeks. What have you been seeing? What's your sort of view on things? Yeah, so this space has exploded over the past four weeks. We can roll it back. I was an early adopter of Poke. So I want to say Poke, which was kind of one of the first consumer agents to really hit the market.
Yes. That launched, I believe it was September 8th. And then I started using it September 11th, three days later. And I remember it so vividly because it was just an unbelievable experience. Their whole gimmick was this negotiation with the agent on how much you're actually going to pay. And it totally blew my mind. And then comes late November and Open Claw gets released. And that kind of takes the actual world by storm of like, there is something coming. And then a few other have popped up here and there.
But most recently, we saw Instinct really blow up the tech Twitter bubble, just absolutely rode a wave, went nuts. And then it was just dominoes. You have Grokbot, you have Muse, you have a bunch of long tail ones like Caddy or Ollie or Pally. And it's really been amazing to see this whole tech bubble is just exploding with interest and an intellectual curiosity for what can these agents do for us on a day to day basis. I think everyone's kind of in the same boat where nobody really knows.
It's the wild, wild west. We're just playing around here. I totally agree. And we should take a moment to talk about Assistant Bench, which is this new really important and discussed benchmark that you're the author of. So let me tee up Assistant Bench a little bit and then I'll ask you a few questions. Assistant Benchmark is a site that essentially compares all of these AI assistants on a use case basis. So we're asking all of these assistants the same prompt, book me a flight to Chicago, find me a vegetarian restaurant within five blocks of my specific location, etc., etc.
And we're comparing the actual outcome on a one shot prompt. Does it actually work? Does it ask follow up questions? How does it perform across 16 different dimensions? How do we compare those all? So we use the word benchmark. Think of it more as like a consumer facing benchmark, not necessarily technical behind the scenes. But if I want to book a flight, which one should I actually use? Who is performing the best? How fast are they actually replying? Got into this whole comparison because I was just doing market research myself, trying to figure out how are these all comparing against one another?
What are the gaps? And published a bit of that. And it was almost two weeks ago, it's 16 days ago today that we launched it. And it's just absolutely exploded. Website has had over 100,000 visitors, have had outreach from basically every single founder. And I think it really just goes to show everyone's curious. People are looking for answers to where this industry is moving. And so talk us through, David, that's incredibly exciting. It's cool, actually. It's funny how benchmarks have become such a staple now of the conversation.
So there's so many launches, it's hard to try everything, which is also an exciting signal. Maybe talk us through everybody who's not instinct and muse. We're definitely going to deep dive on those later. Who else do you think is doing interesting work and what are the kind of directions of specialization that you're seeing? Totally. So we've got a bunch of long tail solutions out there. I think I would categorize it into two groups. You have the B2C consumer facing generalized agents. That's like the instinct or muse.
Some of the longer tail ones, you have Cati, you have Ali, you have Season. There's I believe 64 right now, just in the general category on the site. And then we can break that down even further into a little more specialized within that. So you can have travel specific ones like SOAR or MISO. You can have email specific ones. The other side of the spectrum is going to be more of like your B2B workflow type assistants. And those are anyone from the town to catch to vellum.
And they're all essentially doing the same thing right there. Your assistant, that's actually like more of like an executive assistant, not the consumer facing approach to it. And that's kind of like the general landscape at the moment. So far to my knowledge, there are 122 different ones across all the categories. Incredible. And what about areas of concentration? Where is the most sort of interesting competitive focus? So everyone right now, by and large, is focusing on the general. Horizontal. Yep. Yep. Yep. Just trying to do everything.
And when you think about something like travel, just to pick on that, that's an infrequent high value behavior. Do you think that that category ends up being organized more as a connector into the horizontal agent? Or do you think they have the right to win as a general horizontal agent as well? I think you're going to maybe see potentially one winner as an independent agent. But by and large, I really think it is going to be more of a connector. So I have seven or eight different group chats now with just individuals obsessed with assistance talking.
So there's now over 1200 people across all of the group chats just chatting over and over and over. Travel is the fourth most talked about use case, which is really interesting because when you look at like virality, it tends to be the number one that people are talking about. If it checked into my flight for me, I was able to book my flight. But then you actually look at the day to day conversations of what people are really using it for and most curious about. To your point, people aren't traveling every day.
So it's more of a gimmick, right? It's an attention grab. What are use cases one to three? So number one is more just daily admin type stuff, whether it's cleaning out your inbox, whether it's filing forms for you, the things that are really annoying. And I think that gets to the broader thesis of where I think the winner of this space will go that we can touch on. The second is going to be more agent orchestration, which is really interesting. I think these group chats are a little biased as the third one is development focus.
Okay. These are all group chats of people in tech Twitter, right? So it's not necessarily representative of the world at large. Yeah. But people are talking a ton about agent orchestration. How do you find the most efficient agent or how your agents talk to each other? Development agents. How does this now translate to this coding world on your phone? Whether it be using cloud code from a text component or even take it one step further of, okay, if I'm a designer, can I now make edits to my design files by just like texting my agent of like little small iteration changes?
Yeah. And then the last category being finance. I think that's the most fun one to look into because it is, in my opinion, the least practical There's so many decisions and if you could have an agent make a billion dollars for you, all the hedge funds would be doing it. It's not something that everyone's going to succeed at with their own phone. But it is the most interesting that you give someone, an agent, a powerful tool right in their hands and the first thing they think is, okay, how are you going to make me a million dollars?
Yeah. It's really interesting because one of the critiques, Ben Thompson was on TBPN a few days ago talking about how most consumers aren't looking for productivity in their lives. They're looking to spend time, not save time, which I agree with. But it does feel like finance is an area in which there is a ton of administrative overhead for the average consumer. And if you look at the number of financial markets or financial sort of profit pools that are defined by consumer apathy or being uninformed, that's a lot of profit dollars that could be delivered back to the consumer, right?
So I do think that while I generally agree with Ben that most consumers are looking for entertainment and to spend time, I think there are sort of defining aspects of the U.S. consumer's life that are heavily administrative that these agents can really be a breakthrough for. Totally. And I think that that kind of touches on the general thesis to where I think the winners will sit. And that is that, to Ben's point, I agree. I think the general population does not care about being 10% more efficient.
You hire an assistant to allow it to be invisible, right? The best employee I've ever had are the ones that are just doing their work, getting things done. They update you. Here's what I did. And you're like, wow, amazing. You don't hire someone because you enjoy managing them. And the best, most proactive use cases I'm seeing, I think do fall in the finance category. And I think we're going to see an explosion in a world, though, where it's not cost makers, it's cost savers. And so it's going to be workflows where, like some cool use cases I saw, people are using it to go through their past year of receipts and file HSA reimbursements.
Or one that I do is anytime I book a flight, if the flight price drops, I then have my assistant hit up the airline and get me a travel credit. And I have a friend, this is my favorite use case I think I've seen yet. He hooked up his Grok bot to his sprinkler system at his home and connected it to the weather system. Oh, smart. And now his sprinklers are dependent on the weather and it's dropped his water bill by 50%. Yeah. And so I think these like invisible type agents that are going to be doing these cost saving type workflows fall within that finance category.
But they're really going to lean into these like cost saving activities that create a world that is a lot more consumer aligned. Yeah, it'll be interesting to like, I think the American consumer, nobody wants to hear about saving money. People want to hear about spending money and free money. And I actually think the hook that's going to get a lot of people here is it's going to feel like free money, like going back through and submitting HSA sort of refund requests is going to be free money.
Getting all of your airline tickets where the price has dropped to refund you the Delta is going to be free money. So I think that that's a hugely powerful hook. It's also really, really interesting because this is a sort of infrequent but high value behavior. And that's where my head goes on both travel and finance. Maybe these things will be really powerful, you know, billion dollar businesses that are still vertical connectors rather than general horizontal agents. Totally. It's going to be a very interesting world in terms of the application phase, which I think then, you know, that also something that we talked about last time we went for the walk was, okay, here's what it can do, but now how do you actually interact with that agent?
And it's something that we chat on them, I'm curious your thoughts, right? We both were in the world of iMessage. It's where we live. It's easier for us. But you mentioned your wife and her friends prefer a separate application. And so like, where do you see this world moving? Do you believe that multiple surfaces can exist as a winner or there's going to be one dominantly? Yeah, it's a great question, man. I mean, I think this is why there's so much consumer preference that's going to matter here.
Because I've heard the argument that for folks that are Gen Z, they're more interested in just texting. And for folks that are millennial or Gen X, they're interested in app interfaces. So I think part of it may fragment by generation. Part of it may actually fragment by gender as well, to some extent. I mean, it's a generalization, but the types of things I'm doing versus the types of things my wife is doing, you know, she's big on building a dream board and trying to achieve all of her one year goals and her lifelong goals.
So she really prefers a visual interface. You know, she does a lot of sort of travel is very visual driven for her, whereas for me, it's much more functional. Like, hey, how do I just get from X to Y, ideally with Starlink on my plane, which works better for me? I feel like the iMessage space is more precious. So it feels like a more personal interaction when I'm in that space. And now, of course, that changes if every second thread is with an agent. But for now, I feel like it kind of lives in a privileged place on my phone where, you know, the app on the app grid just doesn't.
What do you think? Yep, I completely agree. I think the personalization that comes with iMessage, just because I already live there, it feels like it's something that's already connected to me. It doesn't feel like an additive workflow. Yeah. And I really have only, though, seen by and large two surfaces appear. It's either the iMessage or a separate application. I've seen a third Signal startup, actually, Sky, they have like an iPhone widget that they're trying to introduce. is a different surface, which is pretty interesting. But then even yesterday, right, we saw Zuckerberg and the Muse team dropped the Muse charm.
Yeah, that's now an entirely different use case for how we interact with the agent. What do you think about hardware? I think it's fascinating. I mean, there's a smart person said this, which is that every person who's working on a piece of hardware thinks that their hardware is the kind of be all end all, like the pendant is the ultimate thing. You know, the watch strap is the ultimate thing. So I don't think that that's right. I think, again, it's going to be very personal preference driven, you know, just as jewelry is like some people like to wear necklaces and some people like to wear earrings.
And I think the same way this kind of form factor is going to manifest is going to be very customer specific. But I think that there's value in having something that is sort of with you and ambient and can capture all the kind of context and requests and, you know, promises and commitments and everything that happens through the day. I mean, I think we're going to have to sort of figure it out together in terms of privacy expectations and society. You know, you go to restaurants in Washington and they take away your phone.
Maybe they're going to be taking away your phone and your pendant. Right. But I think that there's actually something there with hardware. And I just like that we're being ambitious in that direction. You know, I think after 10 years of startups being discouraged from working on hardware, now we've got a lot of interesting new consumer efforts. Right. I think it's going to be super interesting. I have a hot take thesis that this Muse charm is actually less about trying to win the hardware game. And it's more about data collection in the real world for the meta team in terms of like how the real world actually functions, given it has a camera and multiple microphones.
And it's primarily I don't really see it being a mass adoption thing. I think it's going to get adopted by the Twitter bubble and they're going to use it everywhere. And it's just going to feed meta with a ridiculous amount of data to fuel Zuckerberg's future metaverse of mapping out the actual world. But don't you think that they already get that from the glasses? I think they they do to some extent. But I don't know if I mean, correct me if I'm wrong, but are the glasses always going to be on in the same way that this charm will be?
I don't think so. Yeah, it's a great point. I think the charm is meant to be more as you know, as you mentioned, it's an ambient 24-7 companion versus the glasses are more action oriented. It's going to be fascinating. I thought the you know, I've got to try the VR release. That was interesting. But I actually think that the audio only glasses might be a sleeper hit. Yeah, I think there's something really compelling to that. I think people, you know, I mean, the vision of sort of Jane from Ender's Game was all about this sort of all knowing audio only companion.
And I don't know. You know, I think there's a set of people that love to wear that kind of camera on their face. But others may be aware of the fact that it could make people uncomfortable and they're like feeling like they have the technology connectivity without any of the social awkwardness of the camera. Audio. Is incredible. I'm quite bullish on voice and audio. Chachapiti voice yesterday dropped their Chachapiti voice to allow connectors and so you can connect to like your Gmail and calendar. I played around with this morning and to me, it was my second magic moment of this whole AI assistant space.
Really? I was biking to work. Hands are busy. I want to get stuff done. And I just start talking to Chachapiti voice. I just open up a session, start chatting with it. And throughout my 30 minute bike ride to work, categorize my emails, submitted them to different labels. Yeah. Reply to different emails. Sent out calendar invites. Got to the desk inbox zero. It was it was so electric. It's incredible. So this is Chachapiti voice, not Muse? Correct. Chachapiti voice. Amazing. OK, I misunderstood you. I thought you'd said Muse voice.
I totally agree. I think the Chachapiti, it's based on the GPT live one model. It's full duplex voice and it can see all your threads. So you can ask it things like, hey, what are those three projects I started? And, you know, where are we and what am I blocked on? I think it's an extraordinary capability. And it's one that neither, you know, Muse nor Instinct has today, though, though I think Zuck actually did announce full duplex voice. So maybe it's coming to Muse soon. Definitely is coming.
But I do think that for the general consumer, an assistant is most helpful when you are most occupied. Right. It's not like you're going to be sitting down on your couch and spending all of your time using your assistant to do stuff. It's most helpful when you're gardening or cooking or doing things where your hands are busy and you can't actually engage with an interface. And, you know, it goes back to that like invisible assistant where you can just talk and say something and it just goes and runs and does something in the background.
Yeah. And I think that the tech bubble that is chronically online and on their phone and in front of a computer is a very different environment than the general population. And so I'm really excited to see how does this actually hit mass adoption and how does like the average consumer think about an assistant when it's actually like, is it going to be an assistant that solves real pain points for them where they're going to use it when something is really, really painful because they don't want to do it?
Like you get a traffic ticket, you want you have to go pay for that traffic ticket. Great. Just take a picture and let your agent go do it for you. It's going to be interesting to see. So the next thing I wanted to talk to you about was social. You know, we talked in multiplayer generally. You know, we talked about how iMessage is this sort of privileged place to be, which is actually one of the cool things about Instinct. However, in my experience, building agents that you add to group chats somehow feels intrusive or like it diminishes social capital.
Have you seen anything that works well? And what are your kind of theories for how these agents get to a multiplayer state? A lot of people have tried it. I think I've seen three different approaches. The first approach is your traditional attitude group chat, and it's just like the group chat and companion. The second approach I've seen is how Instinct did it, where you have like an agent network where you invite another agent and they're talking behind the scenes on your behalf and they'll update you as to what's going on.
The third way that I've seen it, which I think is the most interesting, it's a new one called Doc by this guy Shane Mack. I think also an A16Z company, XMTP. And they are approaching it more like Granola. So you add an agent to a group chat, but that agent is silent and it's purely a listener. It's taking notes on what people are talking about. And if it has an action item, it independently will ping the person in a separate message thread. And then if you want to see what people are talking about, you can go into the actual docs and see what it's taken notes on.
And that way, it's more of, again, it's your invisible agent. You don't know it's there. It's doing all the ad and work that you would want out of a multiplayer function. And if it needs action items, it doesn't bother the majority of the group. Really, really cool approach. I'm going to try that out. I love that. I love that it's an A16Z company already, of course. You know, my theory on this, if you look at this, and Nir from Oboe said it, and I think he's right, which is the progress the models are making on verifiable domains like coding are, of course, exponential.
But the progress they've made on just prose quality and even like bedside manner has maybe plateaued or maybe is even getting worse. So I actually think that with the models as they are today and perhaps with our own expectations in terms of these social dynamics, I don't think there's a lot of ways that the agents can contribute and enhance the sort of social connectivity of the group. You know, conversely, I think agents that are very utilitarian, you know, one of the experiments I've been doing is adding it.
I'm a collector, you know, I collect records and a bunch of other things. Adding an agent to the chats where I talk to my fellow collector friends and trying to catalog our respective collections and our tastes. And I haven't quite hit it, but it at least feels like it's something that's genuinely additive to the group rather than, you know, this sort of low EQ robot that's chattering when we don't want it to. Yeah, totally. You'll you should you should give Doc a try. I think I think you're going to be quite fascinated by it.
I will. Yeah. Yeah. The other interesting take that I heard this week was that, you know, something about messaging in particular is it doesn't lend itself to these agent interactions because in messaging, it's as if somebody has the mic and it feels uncomfortable to give the agent the mic. Whereas if you look at something like maybe closer to a discord or even a traditional web forum, that's actually something where an agent can make a post and you can either engage with the post or not Reddit style.
But for some reason, it feels less intrusive than group chat. I don't know. Yeah, it's almost it's almost as if we're trying to humanize these agents a little too much. That's right. And I think that comes to another component where if you look at the actual design of the character of the agent, I see some companies that have it try to be like an AI generated human. Yeah. And then you have other companies like Muse that it's like adorable Yeti. And I think it's really interesting to see how the world has fallen in love with this Yeti.
It's just so cute and adorable. Yes. And the like AI generated human. To your point, like I think people feel a little uncomfortable, like it feels weird to humanize these assistants so much in your like in your personal space. So one related question I wanted to ask you is, do you think that the market segments by personality or archetype? You know, we were chatting a bit about this on our walk, but some people prefer a kind of serious, gruff, authoritative doctor. And some people prefer a warm and friendly, you know, sort of a doctor that feels like a peer.
And those are just two different archetypes of doctors. Obviously, no doctor is both of those things are in opposition. Do you think the same will be true of agents? I think everyone for sure prefers a different style of communication. Polk came out of the gates hot, right? The personality of Polk is is quite distinctive. And it's something that I have a lot of friends who are like, I refuse to give Polk any of my personal information because it's so sassy and it actually makes me uncomfortable to engage in a conversation with it.
But I don't think personality is really a defensible moat because it's quite easy to configure. It's like you as a consumer, you text your agent. You say that you would prefer it responds in a certain way. It just saves that to its memory or its soul file. And then you're good to go and it changes. I think we'll get to a place where whatever general agent ends up being the leader, that it'll be quite configurable. And each person can kind of tailor the agent to speak to them how, you know, however makes them most comfortable.
So that's interesting. And I agree with you. But maybe let me try the steel man here for not just personality, but kind of, you know, constitution, for lack of a better word, which is an example of a kind of defining characteristic is presumptuousness. You know, we've talked a little bit about this. And there are some jobs for which you want the agent to be highly presumptuous and you'd rather have it make mistakes 10 percent of the time, as long as it often gets things done. And there are other styles of agents like our finance agent, perhaps, where you want it to be highly cautious.
So I sort of wonder if some of these things, which we're now describing as personality, start to push down into sort of capabilities and, you know, how it tie breaks in a position of ambiguity. And perhaps there's a bit more defensibility around that. I think there is massive defensibility around proactivity. I think that's actually going to be something that separates the winner from the, you know, the meat of the pack. The agent that can be most proactive, that can serve as an invisible assistant and start doing things on your behalf without you even asking.
I think the best, like, gimmick, cool aha moment that we've seen so far is checking into your flight for you. To your point, though, it really does come down to two categories. There's some workflows where you want to retain control and you don't really want an agent to be acting too proactive on your behalf. Yeah. And I think those are workflows that require a change in action on your side. So maybe that is an agent identifies that the insurance you're using is charging too much money and thinks you should switch insurances.
Yeah, it should probably get your permission before it actually makes that switch, because that's that's a direct action on your behalf that impacts you. However, the proactiveness of, hey, I drafted this email for you or I got a flight credit on your behalf. Yeah, there's no action needed on your side. It's just saving you money. It's saving you time. So I think it'll be interesting to see, like, how do we really categorize and break down those those type of behaviors? But it will be a very fine line, because if you if you cross that line once, I think you immediately you lose all trust with your user.
Yeah, it'll be interesting. And we were joking internally, like we're days away from an agent messaging someone saying, I noticed you weren't that into her. So I went ahead and broke up with her for you. Yeah. You know, so I saw that one. It's going to happen, man. But, you know, I think a lot of the magic comes from the presumptuousness. And I think that if you didn't see people posting on X about, oh, the agent did this, the agent did that. It would probably be a sign that the agents are not pushing hard enough in this direction.
Right. I think there's going and actually I'd be curious to know. First of all, I'm, you know, a little suspicious of whether the the Bong Chang post was even real. And for folks who don't know, there's there's a person who tweeted that instinct checked him into a flight and hallucinated his middle name as Bong Chang, and he wasn't able to get on the flight. So I'm not sure if that's just shit posting or if that really happened. But I'd be curious to know if that person is still using instinct.
I'm guessing they are. I would assume so, purely given the traction of that post. They're probably feeding for more and more. If nothing else, at least just a source of comedic popular tweets. Right. Right. Maybe talk, you know, it feels like one of the big things that's happened is that the kind of primitives that were developed during the open claw period have now been productized in a way that the mass market consumer can sort of, you know, digest and appreciate. Do you think there are also poke, which you had mentioned a couple of times, like poke was extraordinary.
I think they did a super clever job. And I think if they had started six months later, they might have ended up being the winner. I think the team is really, really talented there. And I'll be curious to see what they actually have in store for us next, because I'm sure that, you know, that they're a part of cognition. They haven't stopped working on this. I'd be very curious to know if there are capabilities or lessons from the open claw era that you notice that you think are going to tell us what's coming next in consumer agents.
Right now, I honestly see the consumer agent is literally just a replication of open claw, but it's just preconfigured. I don't see much happening in the personal assistant space that you can't do with open claw or Hermes. It's just a matter of fact that you don't have to configure everything and you're not getting error messages popping up every seven hours. And it's it's really easy to use. It's an easy interface. So we'll. we'll see what happens. I think the direction of the space will continue to move towards proactivity.
I think the direction of the space will continue to see more hyper-specialization, which I think goes to your point in your thesis of narrow startups. And how does the space look as we see more of these long-tail opportunities, whether it is, you know, instead of being a generalist, I'm gonna spend all of my time focusing my assistant to be an expert for single mothers with children under the age of three. And it knows everything about how a mother operates her day-to-day life, child development, where it literally feels like that assistant is reading their mind.
And that can maybe be a plug-in into the generalist, but I'm curious, with this being your thesis, how do you think about, I guess first, give us some more insight on what narrow startups are to you and how do you see that playing? Yeah, the theory with narrow startups is that, you know, you can now build software that's incredibly ambitious and incredibly valuable for a very small number of people. And we have this precedent of people paying, you know, 200, 250, $300 a month. So you can build, you know, a $100 billion run rate business with tens of thousands of people, which has never really existed in the past.
A lot of it also is that the capabilities can be specialized to such an astounding degree. And then, as we think about what you're discussing, I totally agree with you. I think there will be this kind of comparative advantage. I think it's gonna be a combination of taste and proprietary knowledge. You know, there may be, just as you, you know, you have a doula that may subscribe to a certain sort of style of, you know, helping women give birth, you may have a personal agent that actually subscribes to a certain school of thought or style or, you know, even a set of lessons that aren't well understood.
And you're gonna hire that agent to sort of provide that service to you. And then proprietary data, you know, you had mentioned potentially having a personal agent just for New York City. Well, I think, you know, just as there might be a concierge that's got A plus taste that's hard to replicate that knows all the best spots that you're gonna love. And there may be another dozen for, you know, another dozen different people or archetypes of people. I do think that there'll be this sort of subjective and objective specialization that happens.
It's gonna be interesting to follow. I'm also curious for your lens, even just noticing our conversation here. We interchange between assistant and agent. And it's been a constant conversation I see people talking about on Twitter. Is the future of this ecosystem, is it gonna be a personal assistant or a personal agent? And what do those mean to you? Oh, that's such a great question, man. I mean, to me, assistant is a sort of lesser form than agent. You know, for me, an agent is somebody, you could imagine them being a peer in a social group.
And maybe we need a better word than agent, but an assistant can only sort of definitionally do what you ask it to do. It can assist you in doing something that you've directed it to do. While an agent has agency, and can just make good things happen in the world. And maybe this parallels a little bit, you know, the sort of agents and AI in the enterprise, where they've sort of started as interns, and through doing work and improved intelligence, they're able to get promoted into doing higher order forms of work in the enterprise.
Like, what does it mean for an assistant and later an agent to be promoted in the consumer's life to do more specialized and delicate work? I'm super aligned. I'm also very interested what ramifications it might have in just like the social layer of humanity, where does the word agent allow us to not personify AI enough? Is assistant, does it add too much human to the experience? And does that make society uncomfortable? And it'll be interesting to see like how the, we're in bubbles, you're in a more bubble, you're in SF, so you're in like the tech hub of the world.
I would say New York is a little behind SF, and then the rest of the world catches up after. So it'll be interesting to see how people actually think about the vernacular of the space. I don't know if it's a bubble, I just think it's like the future, you know? It's like we're in the near future. Yeah. Yeah, I don't know, I actually think that it could increase and improve a lot of social dynamics as well, because you start to think about the agent as a non-human, it can provide a layer of social indirection.
So there might be something that's sort of awkward for me to say to you, or an uncomfortable social dynamic that's occurring in a group, and with a very carefully designed agent, I think the agent can help to, I don't know whether it's provide feedback or sort of organize a conflict in a way that feels non-emotional, because you know that it's not. It's like an arbiter of confrontation. Potentially, yeah, yeah. Now, I'm not advocating for a more passive-aggressive sort of social system, but I do think it's interesting that they can be a sort of source and a destination for social indirection, if that makes sense.
Right, and I think that it will have a lot of significantly positive influences in the world. How do you see AI impacting, right? If we look five years down the line, where do these agents, where do these assistants actually take us? How do they make us happier human beings? It's a great question, man. I mean, I do think that those are the problems that we're really after at the root of things. You know, we talked about productivity, we talked about finances and health. There's a lot of mundane things.
You know, I was looking this up yesterday, but 1.5 million people a year in this country get a DUI. And look, the amount of administrative burden that's probably implied by that, and how many of those people have access to attorneys. So I do think there's almost this sort of bureaucratic pressure that's applied to the average consumer. So I think step one is, how do we sort of relieve that bureaucratic pressure and have people be financially optimized, have their health be optimized, and have them have the same access to, whether it's an attorney or a kind of family office that somebody who's wealthy might have.
Then I think the bigger question, you know, I grew up with my mother saying the hardest part about getting what you want is knowing what you want. It's like, what are we all truly after? And what are the ways that these assistants can sort of engage us? And I don't think it's as simplistic as just goals. I don't know how many people even have goals. I think a lot of it is people are just trying to, you know, find out who they are and turn into the best, most authentic versions of themself.
And this kind of interplay between our own self-development and our subjective experience in the world and these machines could be a very beautiful future for us. Totally. I see this future where I envision this world where my future kids are going to be looking at old videos of the current generation walking down the street, you know, necks craned down, staring at your phone. I mean, what are you doing? It is a beautiful world out there. Why are you spending all of your time staring at a screen?
Yeah, I love that. Because hopefully it's gonna allow us to automate so many of those things. So as a parent, you can actually spend more time with your kids. As an adult in your 20s, you can actually spend more time talking to friends. And I'm really curious to see how it will just change life for the better in so many ways of reducing the things that you just really hate to do, but it's like the life admin work. Yeah, that's right. And that's also why I think that so many times coding agents are not competing with human programmers because they're doing things that no human programmers are doing.
You know, whether it's building this ridiculous side project. I've been building a virtual record shop that you can walk into. I'll send you a link to it after. It's awesome. I would never hire a programmer to work on that five or seven, 10 years ago. I would have never spent the year of my life that it would have taken. So all of this stuff is the sort of things that never would have happened otherwise. And I think that there's a lot of versions of this technology in our life where we're just getting to explore things we never would have explored anyway.
It's not about cost reduction. It's about possibility expansion. Yeah. Makes life more fun, makes life more interesting. You know, to kind of come back to the near term, a thing that I've been thinking a bunch about but I'm curious for your take on is, you know, the lessons of Multbook and what are these sort of agent native systems that are gonna need to be built? And maybe this is some of what Instinct is doing with their kind of instinct to instinct network but it does feel like there's going to be infrastructure and products that exist only for agent to agent interactions that doesn't exist today.
What do you think? It's gonna be a whole new paradigm. I think it opens up, and I think that can be a 10 hour conversation on its own. You have agent emails, you have agent phone numbers, you have what does the world of service look like when it's no longer human booking to human and no longer agent booking from human but it's now agent booking from agent and how does that redefine the entire service industry as a whole? And then you have like the whole cybersecurity component of all this stuff, which is, okay, well, you know, internet came first, then came cybersecurity, agents came first, now comes the security component of these agents and what industry is that gonna create?
The beautiful part of it, although I think that's the beauty of capitalism, there's like so much opportunity and so many problems to solve with these agents and I think it is inevitable that we're gonna see so many startups popping up, trying to figure out all of these niche problems to really build this like new wave of internet and digital connectivity. Yeah, it's gonna be cool to see, because especially commerce, I always think of the web as kind of, there's three webs within the web, you know, just like there's two wolves inside me.
The three webs are the commerce web, the application web and the content web has been in trouble for a while. I think that the commerce web has actually done really well and then the application web has definitely been on a huge upswing ever since coding agents were released. But the commerce web is fascinating. You know, you saw this week how Shopify embraced Muse and Amazon blocked Muse. I think part of the calculus is who's gonna be a winner or a loser when you actually have a new front door for how traffic flows to these marketplaces.
I mean, do you have a view on what specifically happened with Amazon and Shopify and maybe more broadly who the winners and losers are likely to be? Yeah, so I think, you know, first off, Shopify and Amazon are two completely different business models. You have Shopify that's hyper-focused on more democratizing commerce. You have Amazon that makes all of their money on ad revenue and so the whole ad structure of commerce is gonna be really fascinating to watch because when you remove human eyeballs, it defeats the entire purpose of the system.
And so I think that's where Amazon was coming from. Like if you let Muse and these agents start to purchase without them getting any ad revenue out of that, it really takes a lot of their profit out of the experience. So I think there's likely gonna have to be some new paradigm for how they're able to profit in this world where agents are purchasing on their behalf versus a Shopify that is empowering the everyday individual to sell as much as possible and agents democratize this space even further.
So it's just unbelievably advantageous. I think it's gonna be even more interesting when we look at industries that have yet to be touched and that's something we chatted about last week was restaurants where if you look at the booking space, right, like you go to make a reservation, I think we saw Resi started shutting people's accounts down because they were using browser use and they didn't want bots. So it defeats the whole business of if you just let agents take these reservations instantaneously, then it kind of defeats the purpose of the platform of this like human-oriented booking experience.
And so what happens though when everyone has an agent, now everyone's on the same playing field, everyone can go to book the restaurant at the same time, do we enter a world where one, power goes back to the restaurants and they can maybe choose who they wanna give the reservation to based off of loyalty or average spend or whatever it be? Or two, do we go into this world of bidding where everything becomes a bidding war between these agents? Curious your thoughts on kind of what that future looks like.
Yeah, I think you're super right. I mean, very interesting because I think there'll be things that are supply constrained like hot restaurant tables, things that are demand constrained like people that wanna fly to SF tomorrow. I think for restaurants, if you think about with perfect information, who do they want, to your point, it's the highest AOV, highest LTV, which means AOV over multiple visits, plus they actually wanna ensure that they're at near 100% capacity as much as possible. So that's actually a more sophisticated calculation and I don't quite know how they're gonna do it, but I'm sure that proprietary supply is gonna become more and more important versus relying on kind of friction and motivation from the end consumer to act as something that sort of prevents this DDoS.
So I think supply constrained sort of aspects of commerce will be really interesting and then demand constrained will be completely different where you almost have a group buying experience where you have a set of buyers that are willing to buy a product or a service or an experience and then you have the supply side, the merchants that could potentially fulfill that need, come and bid to see who can do it the best experience, maybe at the lowest price. So we've always had this kind of aggregation of demand, but we've never really had aggregation of supply in the same way, especially not in an auction system to benefit the consumer.
So I think that's really interesting. You know, in the Amazon Shopify thing, I think that the Amazon, Amazon probably is rightly taking the bet that they don't wanna be disintermediated and that they'll lose not just advertising revenue, but impulse shopping. I'm kind of curious to know if the agents themselves will be impulse shoppers to some extent, because I think the point around proactivity means the agents are gonna try to make delightful guesses and you know, what is more delightful than a surprise or a gift? So I'm guessing that one of the really interesting sort of dynamics we're going to play out is that these agents are not perfect reflections of their principles.
They're not perfectly rational. They're going to be irrational and non-rational in completely new ways. And I don't know that things like marketing, things like impulse purchases, things like advertising won't matter. I just think they need to be transformed. Totally. I think that opens up a whole new conversation as well, which is what is the recommendation engine of this agentic commerce space actually look like? because AI and LLMs feels like it's kind of broke the internet in terms of quality search. But then if you use an agent and you say, hey, buy me some white socks, there's like 20,000 different types of white socks out there and it's going to send you it doesn't actually know what exact type of white socks you prefer.
Yeah. And so I think it'll be really interesting. I think Muse has probably a head start on this and you know, this is whole Zuck's whole thesis of attacks on all commerce, where they have your Instagram data, they have your Facebook data. So they probably have a little better understanding of what you actually prefer. But it'll be really interesting to see what this recommendation layer looks like in this world of commerce. Well, I think it could be fun for the long tail, you know, because you can imagine to your white socks example, like I could go to Amazon and get white socks, you know, but maybe my agent comes back and says, hey, I found someone awesome on Etsy that can knit you the socks, you know, grandma in Nebraska.
And she just knits these amazing white socks. It's going to take three weeks. Are you down? I'm like, sure, I'm down. That sounds interesting. So yeah, I kind of wonder if it brings more serendipity in because of course, you know, it can examine all the options and perfectly tailor these things to you with no adverse incentives. Right. I would love that. I think that'd be so fun. It would definitely add a little flavor to the mundane life of shopping today. Exactly. I mean, from my point of view as someone who, you know, doesn't take joy in sock shopping.
Well, and look, you can even imagine a world to get a little more futuristic where it starts to now catalyze things in the real world. So maybe instead of grandma who knits white socks and sells them on Etsy, you know, maybe it messages the agent of a grandma and says, hey, would you be up for doing this? So I think this is where the kind of agent to agent spaces get really interesting, where they're all acting on behalf of their humans, you know, some to provide products, services, knowledge and some to make money.
And you know, maybe there's this sort of secondary economy that's much more dynamic, much more high frequency and just much more interesting than the one we have today. Right. I'm just imagining now a grandma in Nebraska receiving a text message saying, hey, this guy in SF wants you to knit white socks for him. Why not? Do you want to do it? Why not? Yeah. Yeah. Why not? Right. But also it sort of sets the incentives in the right way, because if the agents are doing the right thing, they're choosing great products based on product quality and they can assess product quality for themselves rather than, you know, the thing that has the best marketing or the best brand.
And by the way, if they are choosing on a brand basis, it's at least with the explicit sign off of their user. Like I just want to have Nike's because I love Nike is not because, you know, there's sort of this cognitive dissonance that's happening in the marketing and the consumer is actually doing something that's not optimal for themselves. Right. It opens the door to a truly consumer aligned Internet, which I think we have been lacking for quite some time. I agree, man. We touched on the economics of the agents of the one hundred and twenty two that I have looked through.
So I personally tested twenty six of them. OK. Of the one twenty two. Sixty five are paid interest. Thirty five of them are fully paid. Thirty are a freemium model. So once you hit a limit, then you have to pay. Only 13 of them are actually free. When you have Muse and Instinct that are both free, that are also the dominant ones, and then you have open AI that supposedly is going to be releasing something cool maybe next week. Who knows? Mm hmm. How does a long tail startup actually win this space?
And why are they charging money when the big dogs are giving it to you for free? Man, it's going to be really challenging. You know, by our estimate, it's costing something like twenty dollars per user per day to do this in a really ambitious way, you know, potentially hundreds of millions a year for a startup. I mean, I think the two positive trends to bet on are that browser use should get should be deflationary, get exponentially cheaper very rapidly. Now, if you're a startup that is raising today, that's betting on that happening in six months.
Maybe it's a little bit of a game of chicken, but we have seen kind of token prices depressed dramatically and sort of emerging capabilities depressed. So the price of them. So I think it's very possible that browser use gets cheap enough that more more products can offer this for free. I guess the flip, though, is that I don't love the idea of, you know, a customer subsidy being the primary value prop. So I'd love to see the agent that costs a thousand dollars a month that the consumer is so excited about that they're willing to pay like people spend a thousand dollars a month on a lot of different things that are, you know, not necessarily necessities.
They're sort of desires. So what is the agent that is exciting enough that you're willing to pay that kind of price for it? And then that agent will benefit from both extreme market fit as well as sort of deflating costs. Right, right. I think it'll be interesting to see. Super interesting. Well, David from Assistant Bench, make sure everybody checks it out. Thank you so much for joining us. We're going to do this frequently. I mean, I feel like we could do this every week, dude, like things are happening every single day.
So please follow him on X. Follow me as well. We'll all keep up to speed and sort of learn together in the X community and sort of happy day to you and your agents. Appreciate it. Same to you. This is fun. All right, brother. Thank you. Awesome. Thanks for listening to this episode of the A16Z podcast. If you like this episode, be sure to like, comment, subscribe, leave us a rating or review and share it with your friends and family. For more episodes, go to YouTube, Apple Podcasts and Spotify.
Follow us on X and A16Z and subscribe to our sub stack at A16Z.substack.com. Thanks again for listening and I'll see you in the next episode. As a reminder, the content here is for informational purposes only. Should not be taken as legal business, tax or investment advice or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments, please see A16Z.com forward slash disclosures.