How Jev Turns AI Into Software That Gets Things Done
Audit one repetitive workflow you currently handle through forms, routing rules, or manual review. Write down the intent behind each decision—not just the inputs—and identify where a probabilistic classification could safely choose the next action with a human fallback. Start with a low-risk, high-v
42mKey Takeaway
Audit one repetitive workflow you currently handle through forms, routing rules, or manual review. Write down the intent behind each decision—not just the inputs—and identify where a probabilistic classification could safely choose the next action with a human fallback. Start with a low-risk, high-volume task, measure whether the AI is consistently “smart every time,” and only then expand its permissions. The goal is not to generate software faster; it is to make the software you already use capable of doing more work.
Episode Overview
Ben Horowitz and Martin Casado speak with TypeSafe AI founder Diogo Almeida about Jev, a new AI primitive intended to embed intelligence inside software rather than merely accelerate code generation. Almeida argues that the meaningful benchmark for AI is reliable automation of real work, enabled by natural-language intent, program state, and confidence-aware decisions. The conversation explores reliability, probabilistic programming, and why AI may strengthen SaaS applications rather than replace them.
Key Insights
Build smarter software, not just faster code
Coding agents can accelerate implementation, but Almeida argues they generally produce the same kind of software humans already write. Jev’s ambition is to give developers an intelligence layer inside programs, allowing software to interpret intent and make useful decisions that traditional deterministic logic cannot.
Reliability is the unlock for automation
An impressive demo is not enough for a system that runs in the background, touches resources, or becomes a dependency for other software. Almeida emphasizes robustness: the system does not need identical outputs every time, but it needs to be intelligently and predictably useful every time.
Use AI where its strengths overlap with valuable code
TypeSafe’s design principle is to operate in the overlap between what AI does well and what is valuable in software. Almeida specifically contrasts useful classification and decision tasks with areas where models are weak, such as extrapolating precise floats.
Measure AI by work automated, not answers generated
The episode challenges benchmarks based solely on human judgments of model responses. A more practical test is whether AI can reliably automate economically valuable tasks that still require repetitive human effort, such as operational workflows and customer support.
SaaS can become more valuable through embedded intelligence
Rather than making software companies obsolete, AI can help established applications automate the workflows they already understand deeply. SaaS companies possess customer context, workflow knowledge, and distribution—assets that can make AI-enabled features materially more useful than a generic chatbot.
Notable Quotes
"What I want instead is smart software. Instead of automating software engineering, I want to expand what software itself can do, such that things that should be automatable can then be automatable."
"On day one, I draw the Venn diagram of what AI is good at, what is valuable in code. We're in the middle."
"I want them to work in the background, such that someone would trust that to run and not page them."
"I think that SaaS will be one of the largest winners of the whole AI game."
"The highest honor of reliability will be to get to the point when people can program against Jev without making example queries."
Action Items
-
1
Find a bounded automation candidate
Choose one recurring workflow with high volume and low downside—such as triaging requests, classifying feedback, or routing internal tickets. Avoid edge-case-heavy or irreversible decisions for your first experiment.
-
2
Define intent and acceptable outcomes
Write the decision in plain language, list the relevant program state, and specify the allowed outputs. Include a confidence threshold and a human-review path for uncertain cases.
-
3
Test for robustness, not one-off accuracy
Run varied but functionally similar examples through the workflow, including noisy inputs and irrelevant details. Check whether the system remains sensibly useful rather than merely returning the same wording.
-
4
Measure automation ROI before expanding scope
Track time saved, exception rate, human overrides, and downstream errors. Increase the system’s autonomy only after it performs reliably enough that users trust it to run without constant supervision.
Full Transcript
Transcript of How Jev Turns AI Into Software That Gets Things Done from A16Z. Auto-generated from episode audio; may contain minor errors.
Where the fuck is all the automation? AI is so unbelievably smart, and yet it's so useless at all other stuff. It doesn't matter how much AI coding agents you use, the software actually isn't getting better. Maybe you're writing it faster. It's arguably getting worse. OpenAI has been trying to automate customer service since 2020. What I want instead is smart software. I want to expand what software itself can do, such that things that should be automatable can then be automatable. My favorite thing that you guys say is, we built prod, not God.
So good. Because if we had any other big lab leader, even if they had joy, they would cover it up. And then your view is so different. You're like, no, we're going to create a way better world. For nuanced reasons, I don't think we are on the path of RSI in the SaaSpocalypse story. AI can write software, but what if the bigger opportunity is building intelligence inside the software itself? In this episode, Ben Horowitz and Martin Casado sit down with TypeSafe AI founder Diogo Almeida to talk about Jev and a different vision for how AI changes computing.
Diogo argues that coding agents make it faster to produce the same kind of software we already have. Jev is aimed at expanding what software itself can do, giving developers a new primitive for turning natural language intent into decisions that programs can actually use. They get into why reliability matters more than impressive demos, what a new era of probabilistic programming could look like, and why AI might make existing software dramatically more useful rather than simply replacing it. And beneath all of this is the question that drove Diogo to build TypeSafe in the first place.
If AI is already this smart, where is all the automation? Today, we have the founder and leader of TypeSafe, Diogo, with us, who is a bit of a hero to both Martin and me. He is not only building a really interesting product, but creating what we think is a very important movement, so we're super excited about today. Welcome, Jev. Thank you. Maybe you can give us a brief on what is Jev, what is TypeSafe, why is it important? Is this curse-friendly? Yeah. Oh, okay. What the fuck are you talking about?
Okay, okay, okay. Cool. I was actually asked for an elevator pitch, which I tend to ramble on and I don't do well, but I realized my favorite elevator pitch for Jev is where the fuck is all the automation? This is so unbelievably tragic, so much intelligence, AI is so unbelievably smart, and yet, not that I hate on chatbots or coding agents, I love them myself, but it's so useless at all other stuff, and it's tragic. It's tragic so that we have so much diamond in the rough, but not polished for work.
But TypeSafe is making AI for software. We want to make AI powerful, not just for humans in the loop, but to actually build real software, and Jev, to us, is our first model in this whole space to make it way, way better, to make automation. Yeah, and so it's been interesting because it's kind of caught fire in software world. So one of the things that made us go, what the hell is going on here, is every developer we know is calling us and going, oh, this is freaking awesome.
It's great. It's fast, it's great, everything's better. Then how does that, because everybody thinks of, well, we've got cloud code, we've got codex, don't we already have that? What's the difference? And then how does that lead to real automation? Ooh, I wish I had some slopped visuals because I have a favorite slopped visual for this. So I like cloud code and codex. I love the description from Gary Tannenbaum. It's just-in-time software. Incredible way to describe what they're doing. It makes software on the fly, and you can program software in natural language, but it has the same expressive power as software.
What I want instead is smart software. Instead of automating software engineering, I want to expand what software itself can do, such that things that should be automatable can then be automatable. And in a more flowery language, I want to express things like intent. I want to expand the vocabulary of what we can do, and I can talk about all sorts of weird sci-fi things I want, but programming is hyper-specifying valuable things and then infinitely replicating them. It's so freaking cool. And I want to just make that more.
Interesting. So one way to think about it is instead of a tool that somewhat replaces a software engineer with a faster, maybe not even as good software engineer, what you're saying is, no, no, no, we're going to super-empower the software engineers we have to write way, way better, more interesting things. Yeah, yeah. By the way, I just think so many people miss this point, and it's such a subtle point, and it's so important to actually tease it out, which is if you use something like Cloud Coder Codex, which is great, or Cursor, which is great, they write code.
But that code is the same thing a human being would have write. Maybe it's better, maybe it's worse, but it's basically still code, just like code looked 10 years ago. And the thing with Jev is whether or not you're Cloud Coder or a human, you have this new primitive, this new thing that you stick in your code that actually expands the power of software. So instead of writing code, it is something that you include in your code. Which, by the way, is interesting, because it's this very powerful primitive, which would be great if you explain, but it's also a little bit different than how programmers think.
For example, it has this notion of probabilities, or... Also an intelligent layer inside the software. Yeah, think of like a library that you can use natural language to describe what you want, and you give it kind of a state machine, and then it will choose what to do with some confidence levels, which we kind of haven't really had before. Ooh, there's a lot of tricks there. I will jump into one thing first, which is I love the first thing you said, in the direction of where the fuck is all the automation.
I love software so much. I wish I could be writing it all day. I would not recommend being a CEO to people, but whatever. And also, it's wild that AI is so cool, and software has been unchanged in 10 years. Yeah, exactly. That, to me, no one can square this together, and the most we can do is add a little chatbot in the side sometimes that can take actions, but not all actions, because some of the actions are not reliable enough. I just want to give that tiny aside.
I love the point. I'm going to jump back to the point about this is a little bit of a different way to think about it. Yes, I think that machine native doesn't exactly match bits perfectly, and that's actually the art form that we are trying to do in our onboarding. On day one, I draw the Venn diagram of what AI is good at, what is valuable in code. We're in the middle. So we don't output extrapolated floats, for example, because AI is just bad at that.
But things like probabilities are not exactly novel, and it's similar to the is Jev just a classifier argument. Jev is absolutely a classifier. Classifiers are sick. Classifiers, they're designed to be useful. They're designed to be useful, and actually, it's the same interface as some of those ML concepts, because these came from practical people who are trying to make systems work. And what I'm seeing is happening now is that Jev actually, my guess, is that Jev probably is better than having an MLE team from 2019 making the stuff for you, and you can just program it on the fly.
Who knows what could be built? Because there were not that many good MLE teams in 2019 to build narrow things and to be able to collect data sets and measure it and all that, and it is just the beginning. By the way, to this point, do you think there's a slider bar here where on one end is language in, language out, like we have today, on the other end is an existing imperative program, and then you can move between the two, or do you think this is the point in the design space, which is language in, state machine out, which is going to solidify as a general purpose thing for programmers?
Ooh, that's a tricky one. So I will say the answer in my heart. The answer in my heart is that it is a slider. And actually, when I design for the properties we have, I might have made mistakes due to my personal preferences, but intelligence per dollar is my North Star right now. And it could be wrong, just to be clear. Intelligence per second might be more valuable in the short term, but even our interface, like calling the input state, this is intentional. Oh, that's great.
I didn't catch that. It's meant to be the inside of programs. Oh, I love that. So in my heart, we are really optimizing. A lot of the work I do is for even more complicated arrangements of the internals of program state. Can you put intelligence in there? I think this is going to be an ever-present battle to have. We're very intentional about our design. And also, pragmatically, I think certain things happen, like it's easier to make an AI at these milliseconds, so it'll be more like a database for a while than like a standard library thing.
But I would love it to be a standard library thing, too. Can I just pull back for a moment? What is the alchemy that creates a Diogo? I mean, like you speak... When a man and a woman love each other... You speak like an AI researcher, you speak like a systems person, and you speak like a programmer. And normally, these things have been not super overlapping. And you're taking AI, which we've been pushing towards being a being, and you're making it a programmer's tool. So maybe a little bit about your personal journey that...
My history into AI is somewhat unorthodox. I was a mathlete. I was an award-winning mathlete. The way I describe it is, I was good enough at... This is a cringe. I was good enough at math to get girls. So that's quite good. I didn't know that was a thing. You have to get quite good. And what kind of girls do you get when you're that good at math? That's an excellent... Our audience needs to know. Oh, no. We have to inspire the youth here. Don't do it.
Youth, don't do it. It's not worth it. Just be cool and chill and interesting. And don't overcompensate. Wow, I can't believe I said that. So I was a mathlete. But I actually never... Oh, man, this also is a little cringe. I never really liked math. I never really tried. I was just like big fish in a little pond. And to me, math was always the path I was set on, but I hated it because it was always about winning competitions. But then computer science is actually a lot like math.
It's basically like math, but cool and useful and fun and interesting. And I still love giving algorithms interviews. Is it the best thing for me to do? I don't know. But do I love it? Yes. And does it allow me to suss people out really well? Yes, it does. So I love computer science. I consider myself to be a computer scientist much more before AI researcher, despite my history. And what actually got me into it was I also won a Kaggle competition, not from sophisticated math, but from just automating the fuck out of it.
You know, more nested loops. I'd solve it like a systems problem. So that eventually got me... I was forced to speak at NeurIPS, normally an honor, but I hated it because I just wanted to be in the minds. Was that from the Kaggle thing? Yes. Oh, wow. Actually, the Kaggle host of it was Isabelle Guillon, who was the co-inventor of the SVM. Actually, I think the first author of SVM. I'm not 100% sure I'm the first author. And she just basically saw that I was this person who really didn't fit into the research community and then adopted me.
And showed me... It got me to meet all the AI people and my career was just pushed into that direction. And from there, OpenAI? No, it was like a startup with Jeremy Howard. No kidding! Yes. I love Jeremy. Fantastic. Cool. And then Google Brain for a while. Wow. And then Retire for a while. And then eventually I was just kind of tired of not doing anything. And I was like, you know what? Actually, AI is pretty damn fun. And I joined OpenAI because of that reason.
And it worked out really well. Amazing. Really, really well. Incredible. So you said something there that is so unusual in today's world, which is AI is really, really fun. And then the company has such a different demeanor and view of AI than everybody else. And my favorite thing that you guys say is we build prod, not God. So good. Because if we had any other kind of like big lab leader, they'd be like trying to... Even if they had joy, they would cover that. And then your view is so different.
You're like, no, we're going to create a way better world. And it's going to be awesome. And there's going to be not only are there not going to be less jobs, there'll be more jobs, there'll be way better jobs, and everybody's going to have a great time. And just being around you, you clearly believe that. So tell us about that. Because for us, TypeSafe, Jev, it's more than a company. It's a whole movement towards a positive future that most people in the AI world kind of don't like.
Yes. Or they're not with it. I think they don't get it. Yes. It's just a classifier complaint. It's like an ML-level concern while everyone else is having a Jev party. Because it's like, holy shit, we can do all the things that we wanted to do. And I think if you don't get developers, it'll be hard to understand what's really going on. So 100%, I agree with that. I do think that there's a pretty negative world painted that I obviously disagree with. I think it really comes from this monomodal Kool-Aid that everyone believes.
Right, one big brain to rule them all. That's one way to sound much more ominous. Yes. That's what people hear. For sure, that's what people hear. But is that one brain really on the path to rule us all? We have not automated really basic things that I don't think we want people to be doing. There's lots of really, really basic stuff. And I think that, oh man, it pains me when the world is discordant with the reality. And part of the pain is on where the fuck is all the automation.
How can we have AI be so freaking smart? And there's so much financial incentive to automate stuff. Yeah, you could make an excuse for diffusion. I don't buy it at all. I shouldn't name names, but that obviously is not true. Part of the problem is the discordance with the reality and the fact that AI has so much potential. is what made it really tragic for me that we had not released this. So now it's like a little bit of a party for me, but I was afraid of AI.
And all Jev users are like... There's the happy AI, the people on Jev, and there's the morose AI, the people who are not. It's really... It's quite a kind of fascinating dichotomy. It is really... Well, I'll give you... And to your automation point, I had a funny conversation this morning with David George, who runs our growth fund, because we were talking about the new tools. I was like, have you tried the Muse thing? He's like, oh, it's awesome. I was like, what did you do with it?
He said, I finally canceled my New York Times subscription. And I was like, that is hard to do. But, you know, it's a kind of a... It's a very tip of the iceberg of the things that are horrible things to do that we need to automate. I think that if we were going to be really intellectually honest, and we are really aiming for the North Star of automation, we cannot fall into the same anti-patterns that AI has fallen into, which is really focusing on outliers and demos.
A lot of people ask me, what are your favorite use cases? And I'm like, I'm not sure if they work. I want them to work in the background, such that someone would trust that to run and not page them. And people can build on top of that too. Composable, composable. Composable, but like other things like safe, right? Like it's a different type of safety, where if you want it to actually run with resources associated with it, with access to things, you need guarantees for that, or at least statistical guarantees.
And so it doesn't go rogue, breaking the hugging face, that type of thing. Well, I don't think our models will be doing that anytime soon, unless someone does the software to do that, which would be very cool flex, very cool flex. I should figure out how to give credits for that. But not in a way that we're not responsible. Right, right, right. I'm just curious, how long has this intuition been percolating? Because I remember talking to you maybe in, was it 2017? We did talk about that, yeah.
Yeah, and then a lot of these ideas were in, you were talking about data being important, and you were talking about you want to focus on the task. So I just like, was this, did you know that this was going to end up being a classifier, or was this just an intuition that there's just kind of another way to view this entire kind of AI movement? So actually, a fun story about that chat in the talk from 2017, I think my talk was actually on a very similar theme.
I think it was called something like AI modular in theory and flexible in practice, which is very software-assistive. Yeah, totally. So I'm a little bit consistent in that. I think that this really started right before chat GPT. Like right when we released these things, I did not have intuition about this. And honestly, I was not even, I was very, very pleasantly surprised by the generalization capabilities of RLHF. When is this? Must be end of 2021, like fourth quarter of 2021. Like we were, it was really, really general.
Like if you read the paper, it's unlike other papers that are like trying to prove their point. It was us actually, you know, scientific method-ish trying to disprove, like, is it cheating? And, you know, my favorite query was, why is it important to eat socks before meditating? We'd made sure that was not on the internet beforehand. And like the models were able to like make plausible human-looking answers for this. And that to us in the team was the thing that clicked, like this is not cheating, which you should always be afraid of cheating in ML.
And then what really got me burnt was we released it. You know, we did, you know, I'm obviously a big capabilities guy. I did a lot to release that model. I really thought that the model had like a decent chance of being AGI. And when it didn't, that was like when my whole world came crashing down. And I was like, why? So you were kind of on the other train for a bit. Like the crazy train? Well, no, I'm just like, RL generalizes, like maybe we have AGI, like.
RLHF generalizes pretty well. RLVR is the thing that doesn't generalize as well from what I've seen. And AGI in, well, I was just saying more, I mean, like, you know, you were behind chat GPT, you were behind these early GPTs. That was a very different goal, which is like creating a chatbot that will talk to the human being, which is not a programmer's tool, et cetera. So I'm just wondering like. Oh, well, actually early, early, like 2020 OpenAI, when we talked about AGI, people used to describe it as Ilya and every if statement.
So it's not, it's kind of like. But like, is it like part, we were talking about OpenAI culture. Part of it is that it's like intentionally vague. So it's a wide, like tent, so that everyone can be inside of it. But like, I am not, for nuanced reasons, I don't think we are on the path of RSI. And I still don't think we're in the path of RSI, and I did then. I do think that what OpenAI defined as AGI is extremely doable. Automating most of the world's economically valuable work actually sounds like, oh man, I don't, like there's a lot of work out there.
A lot of it is very rote and simple. And like by volume, in order to be able to like outsource work, you need like simple instructions that like basic people can do. And as far as I can tell, the intelligence of that has been available in the models for like quite a while now. And like my, oh man, you know, chip on my shoulder is like, why is this not available? And then since RLHF, the AI industry just like kind of bifurcated into gigantic over-promise, under-deliver.
I think GPT-3 was actually quite calibrated back in that day. But because humans evaluate how good the models are, it looks really good because they are the judge, but we've been optimizing that judge instead of the automation part. And that has been the missing thing. So I would say that it was really, really then that it like hit me, you know, like why is this thing not more useful? And so you think that the measure that we should have is to what extent can you automate actual productive tasks?
I would like that. When you say over-promise and under-deliver, that's the dimension in particular you're talking to. The ability to automate tasks. I like, I think in my heart, it's like cool sci-fi, you know, and I think that, I think that that is the canary in the coal mine for cool sci-fi. Like, are you really telling me that math is solved or like even like two years ago, GPQA, that Google proof question answering is solved, but we still can't handle a drive-through, right? Like, it's a very hard thing to hold in your head at once.
And I think a lot of people don't have good answers to that. Can I just test one thing, which may not make sense, but I want to just, I mean, isn't there an argument though that like the distribution of the real world is different than the digital world, right? It's heavy-tailed. There's a lot of exceptions. We don't have all the data. And I mean, couldn't it be the case that the reason we're not doing productive stuff in the real world is just like, we're not, we don't have the data for that distribution.
We're not training on that distribution. And this is why it's just been basically relegated to like these lower dimensional manifolds, like whatever math or code or... I don't entirely buy the data argument. In my opinion, I do believe that there's a long tail for sure. Like that would be kind of crazy to deny. And I don't think that in my, like canary in the coal mine situation, we need to automate that long tail. Like I think that we need to be incredibly pragmatic on everything and like building reliable software is always an investment, right?
What were the three great virtues of a programmer? Laziness to not to do it again, hubris, and there was a third one. Now I remember, this is from the Pearl days. Yeah, there's a third one. I wish I could remember it. But like, it's about like the laziness to like spend like 10 hours to do like the five minute task instantly and to never have to do it again. Like it only makes, like it should be an ROI decision for people who like automate stuff. Like I would just like it to be automatable.
And I think that people will just make like new kinds of work, hence the Jev in Jevons, new kinds of work once that stuff is doable. But like as like a benchmark, I feel like it's useful to see, can we actually automate the stuff that it really, really looks like, yeah, I should be able to automate. OpenAI has been trying to automate customer service since 2020, you know, like it's, you know, like it's not. It's just pretty amazing. It's wild, you know, it's wild. Well, and inside, I mean, inside companies, there's very little that's automated right now.
Like, and the projects haven't worked. Other than programming has worked amazing. Can you maybe classify the types of problems you think that are easier to automate now? Because it was kind of interesting. So we've actually looked at support before the current generative wave. And it was interesting, you'd meet a company and the company would say, we answered 95% of all like, you know, like help desk calls. I'm like, that is so many. But then you actually look at the data. They're all the same. And you realize it's all password resets.
And then like, but if you did it by like uniqueness, there's only something like 50%. So it just feels like when you're dealing with humans and natural systems, like there's just kind of this very kind of, you know, like a long tail of exceptions. And so to what extent did like probably every hour I have somebody ping me and like, I'm using Java for this new use case. I'm like, I had no idea, you know, like, you know. And so like, to what extent did you even predict like the broad range of use case for it?
Like, did you assume that was going to happen? And have you been surprised by that? Extremely surprised. Did not assume it would happen. This launch was like not something, like if anyone expected this, they are probably insane, right? Like it is, I don't think someone could expect a chat GPT for developers, because chat GPT was for, you know, like, you know, the normal users. And it's weird. I actually don't even know what percentage of the people who are part of the Javaparty are developers themselves. I can't imagine non-developers using it.
I don't know how they would use it. But even my non-developer friends are just like part of the party and Twitter and memeing and everything like that. So number one, phenomenal. Number two, this will be hard to convey in this short message, because like it's been like blood, sweat and tears for years now. Like the amount I care about reliability is, it's a lot. Like reliability is what this thing is. If you don't understand that, it'll be very hard to make like a copycat that's benchmaxed.
Like it's, I feel like every nine of reliability is going to be so valuable for everyone, even if it's not the most valuable thing market cap wise, because it will just enable new applications. And like we are fighting for like all sorts of like weird nines of reliability that like we don't even fully understand because we are just like, you know, like really getting this like electric motor of AI, of intelligence, like into people's like workstations and they can figure out what to do with it.
What does reliability mean in this context? Is this like availability of the model or is it like I call the model and it returns the same thing? Or like, how do I think about reliability? Yeah. For something that's inherently kind of stochastic. Not so much the former thing. And the second thing is closer. Like I would describe the first thing as kind of like uptime or SLAs. The second thing I would maybe call closer to determinism. Something thirdly, I would consider more like robustness. So robustness I would kind of describe as similar intelligence every time.
Oh, interesting. So like not exactly determinism because I think determinism, it's useful for unit tests, but not real systems. Think about like if you add a UUID to a prompt, it should be the same because it's the same functionally, but it's not exactly deterministic. I think that there's another layer of it that I don't really know what it's called yet. Like maybe this is what I would call like some form of intelligence, which is it doesn't have to be the similar function every time, but it needs to be smart every time.
You know, like if you were in that situation, would this be an understandable thing for a human to think? Because a developer can program around that. And actually to me, the highest honor of reliability will be to get to the point when people can program against Jev without making example queries. Like when you just trust it, you'll be in like perma flow state, just creating crazy software. And like that's where a lot of software is today, right? Like I don't think it's totally unrealistic, but I'm going to be fine.
By the way, this is kind of a weird question. So like, I mean, feel free. Like if it's too weird, just feel free. But it occurs to me that actually the value of things like coding agents goes down if you have a primitive like this in a way, which is like you could be like, you know, whatever, some, you know, Codex builds all the software for me, but it doesn't actually use Jev. And so like the software itself that it creates is somewhat limited. Or you can be like, okay, I as a human being, I will write the software without using a coding agent, but I've got this very generalized primitive that makes writing software easier.
So like, do you feel like, see a future where it's like the coding agents using Jev, and then you're telling the coding agents and then do you have like redundancy? Or do you feel it's like humans? This is more of like a coding agent question than it is a Jev question. Oh, yeah, for sure. My vibe is that I'm not in the coding mind as much as I'd like to be. So you are, you two might be in there more than I am, which is sad.
But my experience is that they are really good at syntax and really... They're bad at semantics. I would say incredibly bad at architecture. Yeah. So like to me, architecture is like the most human creative part of software. So I love using coding agents. I think that Jev is almost certainly not in distribution. That would be spooky if they trained on our user data. So it's probably not. But I think that when it is in distribution, I see no problem with like having it to do the syntax.
And the thing with architecture is that maybe the models are actually like not just crap at architecture, but maybe they're 50th percentile at architecture. And if you don't know anything about architecture, it would be fine. So these are all like gray area trade-offs in order for you to navigate. And sometimes speed is the knob for your company or project to turn. Like you're willing to do a 50th, like a 50th percentile architecture instead of a 60th because you want to move faster. of like codex work overnight or something like that.
Actually, Ken, along those lines, one of the interesting things or phenomenons in the market already is that, you know, when the coding agents came out, it was the SaaSpocalypse and all their values dropped through the floor. And then when Jeff came out, every SaaS company is like, this is the greatest thing ever. So explain that. I don't know what else to say, right? Like, I think it's like quite natural. Like in the SaaSpocalypse story, the story that I feel like has panned out really poorly is that software is very cheap and perhaps easy to replicate, which I think I could I could believe the former.
I could not believe the latter because a lot of the stuff happens beneath the hood. I'm maybe overly a software fanboy here. Yeah, all of us. Yeah. Okay. I didn't know where you might be the coding agent. We have a lot of legacy around that. Yeah. Yeah. So I don't think that really panned out. So SaaS seems like maybe the markets don't agree. But I think SaaS is providing the same value it used to. Maybe the markets are just scared. But I think that SaaS will be one of the largest winners of the whole AI game.
And I want to like work really, really well with like all the biggest, most boring, most like in the know of user problem SaaS companies, because I think that they are the best position to know what workflows to automate. What do people need? Like that's what their bread and butter is. And to spend the big like, you know, software is always a capex investment. And like you spend it ahead of time in order to make this experience even better. That gets, you know, like distributed to all of that massive users.
So I think that it's going to be I'm not going to forecast anything about the financial markets. But I think as far as like a capabilities games goes, it's going to be like an inverse SaaSpocalypse. And I am so jazzed about it. I should make a name. Yeah, I got that. Yeah, that should have a name. SaaSapalooza. Oh, that's a little too fun. Well, the SaaS applications are going to all of a sudden get like dramatically more useful. And by the way, you know, the kind of capital investment, like so much of a SaaS company's capital investment is actually getting to all the customers.
And so if you've gotten to all the customers, and then you make, you know, not just put a chatbot on your SaaS product, but actually make the software like way, way better. That's a hell of a thing. I don't know if this is a realistic dream or not. But I think that there's a world where like the multi choice forms just disappear. You know, like I feel like they're like, they're always like something mapping natural language that usually the software already has into like a javelike output.
And I think it's literally, it's literally from the 80s. It's like, it's called we used to call it 4G LST, fourth generation language. Yeah, actually, also, I think this is from the 80s. This might be an insult. I was born then. Like, I think that do what I mean, is going to be like be taken to the absolute next level. If I should shout out one Java application, I don't know if it's reliable. So I can't promise anything. But it was so freaking cool. Someone was using like a voice to control your computer.
And it was basically constantly making decisions on like, is this a command? Or is it inserting text? Where's inserting text? Like, like, that sounds so unbelievably cool. Like, like, I feel like interfaces could just completely change. And maybe we're gonna have to make it like cheaper, faster. Yeah, then you're at Star Trek. Well, there's just such a profound intuition here, which is if you use AI today, to generate software, right, you're still creating the same software that you did before. But if you actually look at like the average PR for a large company, it's like 10 lines, right?
Seriously, no, I've been at Google, we actually did the study. So it's like 10 lines. So they're like, you're automating 10 lines. And by the way, those 10 lines are, are like, you're part of a learning from a customer or something. So like, you've kind of optimized something that's actually pretty minimal. But what it doesn't do is provide new capabilities to the software, right? It's kind of automating this thing, which in the limit ends up being relatively minor. And now like, there's actually a new capability.
And so like, it could just be the case that just software just actually gets better. And by the way, even before Jever, it didn't even occur to me that like, it doesn't matter how much, you know, AI coding agents you use, the software actually isn't getting better. Maybe you're writing it faster. It's like arguably getting worse, just because like there's less oversight. So I think this is often more insecure. For sure, for sure. But you're, but you're actually now can make an argument like, like, like, like apps will have new functionalities as a result of this, because there is this new primitive that you're providing that I mean, like in a way like it speaks natural languages, and it can reason, but it marries that to a state machine.
I, if people take that as a takeaway, that would be like the greatest compliment ever to what we are doing. Like, I actually feel like it's almost too grand of a vision to expand beyond the three logic gates that we have into like, you know, it, our types are kind of like one of the same logic gate, but like one that's like a little brain in there. Like, that would be the greatest compliment to like the type safe legacy, because like that is, that's a very huge, non trivial, huge thing for the world.
I'm not gonna like over promise under deliver that, but I will fight for that. Yeah, I mean, listen, I mean, there's, I think, pretty open questions to what, like, how deep can this get as far as like, like really serious stuff like state consistency or durability or like real systems level stuff where you actually need to like provide strong guarantees. And so 100% this will change things like whatever, analyzing logs, analyzing emails, providing a UI talking to the human like that for sure. But like, you know, you could argue that over time, this becomes like a smart database, you know, and also air traffic control system, which we really need.
A little scary. Like I think automate the easy work before the hardware is always my philosophy. But I also think there's going to be like an entire era of probabilistic programming that's opened up. Like my, by the way, you know, there's a huge history of probabilistic programming. I basically died in like the seventies. I'm familiar with it. I actually think it's going to be like with the same, like, you could also call Jeff like neuro. So your, your co-founder Eric came from that background. He was telling me.
Oh, cool. Oh, yes, yes, yes. He did a lot of biology. It goes up and down. But like what I mean is a more, I'm not a fan. My brand is pragmatism, incredible pragmatism. I'm not a fan of like biologically inspired stuff at all. It's never worked. Have you ever noticed that? I think it's never worked. It's useful to motivate crazy people to work on things for decades until it works. And then they refine it into like the engineering story of AI, neural nets for sure.
Yes. But like, you know, a lot of the stories about how it worked were not accurate. So like the hierarchical features of applications really did end up working because like otherwise ResNets wouldn't have worked. Longer story. I do think that it opens up like, like from a systems perspective, I'm not excited about this part because it's really, I'm excited for the world, not about me programming this because it sounds like really complicated. But I think that as we have like lots of intelligence at lots of like different cost and speed trade-offs, the super systemsy types will be making trade-offs at like, you know, like Jeb's going to be like a thousand times too smart for them.
They just want like an approximate link to have an approximate guess to like optimistically route here and there. It's going to be like so crazy, this type of stuff that's available in the extreme systems. And the good news is we get to like rebuild systems again, which is great, right? We have a new, no seriously, we have a new primitive. It's kind of a new way of thinking about doing software. Like, I mean, we did this, we did this for the internet and we did this being friend to client server.
I mean, we do this periodically. And by the way, just because of the cybersecurity issues, we probably have to rebuild almost all the systems to just make them safe, I would think. I think it's pretty clear that they're like, there's not, or at least a critical, the critical infrastructure for sure. Yeah. Do you think about this more in terms of like apps, SaaS, analytics, or more in terms of like systems, foundations, or all the above? For what I would think of? Yeah, just general application for this.
And when you think about like, like you're working on Jev and like, and you kind of envision that people are adapting it, like, you know, like, do you, maybe do you even have an opinion? I have a little bit and it's, so the way I think of it is a little like, like deep into like the TCP guts, you know, like UDP, TCP, you know, like it's unreliable. You're speaking my language. Exactly. And like, when I think of AI, and this is why I care about intelligence per dollar, to be clear, when I think, and how I got to this conclusion, I work backwards from AI based economic revolution, AI everywhere, you know, sci-fi and everything, like all the software is AI all over the place.
And I asked myself the question, what percentage of the calls to AI, I imagine it's like a function, which, what percentage are like for human consumption, where you need that style? And yeah, and it's going to be like many nines. And actually from that same question, how many will be at the first layer versus like deep in the guts, right? And I think that it's going to be many nines in the guts, but it will start at the first layer. But like, we need to, if you don't aim for the guts, that's weird.
You don't aim for the guts, it's going to take you a while to get there. Right. I think people don't understand to what extent like AI was kind of ships in the night with software. Like, even if you try to embed AI in software, it kind of like didn't behave, right? Because software doesn't really take natural languages, and you like do all this weird stuff, like you stick in the prompt, like, here's the JSON output that you want, and here's a schema, and it would never listen to it.
And so, and so what you ended up doing is just taking the output and giving it to a human, you're like, to hell with it, right? Or another LLM, that is what a while loop is, like the agent while loop, right? It's like, from first principles, it needs to be human in the loop, which is the chat. Or an agent, which is the while loop, because like the natural language needs to be fed back into another. Yes, and I will say, I have watched this happen.
There was almost like this kind of like, five stages of grief. Like, you know, people will pick up AI and like, I'm going to use this, you know, within my software, right? And then, you know, and then it would like go with like, you know, whatever denial, like try to make it work, and like, anger, then they go to acceptance, which is like, okay, never mind, I'm just going to give this to another LLM to a human being. So it's been very shifts in the night.
I think this is the first time I have seen when someone was like, actually, you can take an LLM, you can take AI, and you can actually map it to like, like a state machine. And you can do that productively. And I hope so. I will not want to over promise under deliver as well. Like, I don't know if it's ready for all the applications that have been over promised. I really, really wanted to, and my team will fight for that, obviously. Like, we really, really care about reliability.
We could have released so much sooner. I don't think people realize that. And I don't think that, honestly, I don't think that they will, based on what I see at the Twitter discussion. I think people will never get it. But like, it'll just have like that good vibe of like, oh, I can trust this. It's the anti-frustration machine. It's a, I hope so. I hope, do what I mean, right? Like, to me, that is about like, smoothness in the world, like having everything that just move more smoothly together and interlink like gears.
I actually have my whole like, AI utopia on like different axes that I really, really want. And like, do what I mean is a huge part of this. You know, like, imagine if all technology just did what you mean. That is like, that's not sci-fi. Look how smart AI is, right? Yeah, no, it's a, it's amazing. And maybe that's the thought to close on. Do what I mean. Thank you, Diogo. This has been a great conversation. Thanks for listening to this episode of the A16Z podcast.
If you liked this episode, be sure to like, comment, subscribe, leave us a rating or review, and share it with your friends and family. For more episodes, go to YouTube, Apple Podcasts, and Spotify. Follow us on X, A16Z, and subscribe to our sub stack at a16z.substack.com. Thanks again for listening, and I'll see you in the next episode. As a reminder, the content here is for informational purposes only. Should not be taken as legal business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16Z fund.
Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments, please see a16z.com forward slash disclosures.