Why Specialized AI Could Beat The God Model
Audit one recurring AI workflow today and ask: does it truly require a frontier general-purpose model? If the task has predictable inputs and a bounded output—such as classification, cost estimation, routing, or policy checks—prototype a small specialized model or structured-output layer. Benchmark
48mSummary published by 1% Better, updated .
Key Takeaway
Audit one recurring AI workflow today and ask: does it truly require a frontier general-purpose model? If the task has predictable inputs and a bounded output—such as classification, cost estimation, routing, or policy checks—prototype a small specialized model or structured-output layer. Benchmark it against your current approach on accuracy, cost, latency, and failure modes. Keep a human owner for the decision, rather than outsourcing both execution and accountability to an agent.
Episode Overview
Amjad Massad of Replit and Alex Atallah of OpenRouter discuss why enterprises may resist dependence on a single frontier-model provider. They explore model routing, specialized agents, proprietary data, safety checks, and the case for combining multiple models rather than relying on a universal “god model.” The conversation argues that AI systems will become more useful, economical, and controllable when organizations retain independence and use fit-for-purpose models.
Key Insights
Own your intelligence, not just an API contract
The speakers argue that AI is becoming a permanent strategic capability, not a feature that can be checked off once. Enterprises need internal knowledge of their data, evaluations, model fit, and costs so they can change providers and maintain differentiation.
Use the smallest capable model for the job
Large general-purpose models can be unnecessarily costly and risky for narrow tasks. Specialized classifiers and decision models can produce bounded outputs, lower the room for misbehavior, and be trained on proprietary data that makes them durable.
Specialization restores accountability
Alex argues that a general agent handling every domain can cause people to lose situational understanding without anyone taking responsibility for that loss. Vertically focused agents allow teams to tune oversight, quality checks, permissions, and acceptable failure modes by workflow.
Model diversity is a product advantage
OpenRouter’s thesis is that different models have different strengths, costs, and training histories. Routing or fusing models lets organizations stay closer to the capability-cost frontier while avoiding lock-in to a single vendor.
Treat agent actions as security-sensitive events
The speakers propose using a fast, separate model to inspect tool calls and agent-to-agent messages against policies that the acting agent may not see. This adds a defense layer for data isolation, sandboxing, and alignment with the original task.
Frameworks or Models
Model Routing and Fusion
1. Define the task’s quality, cost, latency, safety, and data-governance requirements. 2. Evaluate multiple model families rather than defaulting to one provider. 3. Route each request to the model best suited to that task, or combine outputs from models with complementary strengths. 4. Continuously benchmark results as model capabilities and prices change.
Independent Agent-Action Alignment Gate
1. Capture the original agent’s system prompt, current context, and proposed tool call or message. 2. Send those inputs to a separate fast decision model with additional policy constraints. 3. Classify the action as allowed, blocked, or requiring escalation. 4. Execute only approved actions and log rejected or uncertain actions for review.
Notable Quotes
"Every company needs some AI practice, AI capability, and that will compound over time."
"The more work you give it to do, the more understanding of what's going on you're sacrificing. No one new is taking responsibility for that sacrificed understanding."
"Most of the times, like a lot of the use cases, even unstructured use case don't need that capable model."
Action Items
-
1
Create a task-to-model inventory
List your AI workflows and record the current model, data accessed, output format, cost per task, latency, and consequences of failure. Mark tasks with bounded inputs and outputs as candidates for specialization.
-
2
Build one narrow classifier
Choose a high-volume decision such as ticket routing, cost estimation, policy classification, or lead qualification. Train or configure a structured-output model with discrete labels, then compare it with a frontier model on a held-out set.
-
3
Add an independent action gate
Before an agent executes sensitive tool calls, send the system prompt, proposed action, and policy rules to a separate fast checker. Block, escalate, or log calls that violate access boundaries or the intended scope.
-
4
Benchmark for independence
Run the same representative tasks across at least two model families. Measure quality, cost, latency, reliability, and data-governance fit, then avoid designing critical workflows around capabilities unique to one provider.
Full Transcript
Transcript of Why Specialized AI Could Beat The God Model from A16Z. Auto-generated from episode audio; may contain minor errors.
You saw the SpaceX S1, and it's like, oh, $30 trillion. It's like, what is the world GDP, $100 trillion? Both Stripe and OpenRouter really want lots of new companies in the world. We don't want everyone to be a part of one giant company. The reason why the OpenAI hacks have been so destructive is because they're so capable. It's like nuking a butterfly. When the models get more intelligent, the risk actually will continue to get higher. And yet, no one knew it was taking responsibility. We're going to slowly realize how good we've had it with deterministic code.
Remember the days when computers did exactly what we told them to do. Are we going to prevent the models from deceiving users during training runs, predictably? Will a model that's big enough and powerful enough suddenly stop deception, stop sandbagging? Combining models with different strengths, rather than locking into a single provider. Amjad explains why enterprises may increasingly want that independence too, owning more of their intelligence and choosing the best model for each job. We also get into specialized agents, model routing and fusion, the trade-offs between capability, cost and safety, and what happens when AI systems start training smaller models to handle specific tasks.
Welcome to Agent Z Podcast. We're here with Amjad of Repl.it and Alex of OpenRouter. This is the first podcast that Alex has done since the acquisition, so we're really excited to have both of you. Thank you. I'm excited to be here. Alex, let's start with that, actually, if you can briefly share. Obviously, massive acquisition. Amjad is an investor. We're also, of course, an investor. The biggest shareholder, but who's counting? Alex, why don't you give us a little bit of the backstory? How does an acquisition like that even happen?
You get a DM from Patrick one day. What can you share? Hey, how much for OpenRouter? I had talked to Will Gabryk, the Stripe president, a long time ago, a couple of years ago when we were doing our Series A, and we just stayed in touch. We had a lot of Stripe work streams going on with various teams at Stripe. There were always things that we were doing with Stripe. We presented at Stripe sessions, and so they always felt close. And then, yeah, just in July, I believe, they reached out, wanted to chat, met both of them in person, and it just progressed from there fairly quickly.
They're very efficient, and they were very founder-friendly about the experience. I was really impressed with the whole thing. Did you want to sell? Did it even cross your mind before they reached out? No, we were not thinking about that at all. I did really respect the company, do really respect it today, and of possible acquisition options for us, it was, I think, my top choice. It was an interesting idea. As we fleshed out the reasons why it would make sense for both companies, it got more and more interesting.
It was really clear how aligned they were with us, having autonomy over the brand and roadmap and product, and keeping OpenRouter doing what it's already doing, just much faster with a much more serious go-to-market plan, and then some better-together stories between the two products and the two companies. And then culturally, there were just, I think, in terms of mission and values and building a neutral, trusted platform that businesses can depend on and scale on top of, that's also really developer-friendly, with the best possible developer experience to encourage new companies to emerge.
That alignment was there. And there was a bigger picture kind of alignment, too, where both Stripe and OpenRouter really want lots of new companies in the world. We don't want everyone to be a part of one giant company. We want to create really good incentives and really easy, streamlined workflows for people to start new companies and grow them successfully and make both lifestyle and venture-backed businesses on top of really good, reliable, and price-efficient infrastructure and a marketplace that works. And I really want that future. And Stripe demonstrated that they've been wanting it and building towards it for many, many, many years.
And in many ways, payments and inference are going to blend together for companies of the future. I'm curious, I understand why Stripe's incentive is to have a much more vibrant startup ecosystem. I understand the kind of moral argument and why you would want that, but why is that good for OpenRouter? Is your model for OpenRouter is that it's a network-effects business? Is it like a network? Well, for us, I think there's a couple different problems OpenRouter solves. Allowing you to build a company that uses AI or augments intelligence with unique data and other services without model lock-in, without vendor lock-in.
Allowing you to be on the Pareto frontier continuously as the ecosystem grows it. And to do that, it's a lot of work because there's all kinds of little lock-in that appears. We also want companies to feel like that they can add more than just prompts on top of a single model. There's a lot more to building unique intelligence. And I think a big component of that is neurodiversity. You really need the power of multiple models that are trained in different ways, including some of your own, to do more than chat GPT or Claude would on the task if someone who's thinking about buying from you as a potential customer is wondering, well, what if I just use the model directly?
How do you really show that you are significantly better and able to build a business that matters? And I think a lot of it will involve neurodiversity and blending powers and good data from multiple models. Another component is helping people get really good cost efficiency. There are a lot of businesses that just cannot, do not emerge until they become cost-effective. And creating an environment where we can help drive down costs by building an efficient market is crucial to making that happen. Otherwise, why lower my prices?
As a provider, we have a captive market. And I think that's a key point of marketplaces that was just totally missing from AI before we showed up. There was just one player, OpenAI. It could have been a very strange world. I'm not saying that we did all the work. Of course not. But helping people choose new models and explore new models and learn what makes a closed-source or open-weight model actually good at your task involves seeing what the whole ecosystem is doing and learning automatically. LLMs are not things where you can just enumerate all the features on a webpage.
It's impossible. You have to see how they're being used to know what they're good at. You started the company three years ago. I'm curious what has surprised you the most about the evolution of open and closed-source models to the present as it relates to model performance or just people's perspective on the variety of models they have available to them. People have been more open-minded than I thought they would be towards open-weight models. Typically, there's a lot of brand, especially in enterprise. Typically, there's a lot of brand trust where enterprises in general are like, I don't really know how to tell the difference between these things, so I'm just going to buy the one that all the other credible enterprises are buying.
And there's a lot of enterprise lock-in with a mentality like that. And that just didn't happen that much. It did happen a bit, but we saw a lot of enterprises want to explore new models. It was very much good for a marketplace. Enterprises wanted to diversify outside of just the proprietary frontier model labs, both for cost reasons and for differentiation reasons. They wanted to own their intelligence so they could, I think, one, keep their talent, have an internal AI practice. AI is just a huge strategy topic.
It's not like you go to your board and you're like, oh yeah, we fixed the AI problem. Quarter complete. Your board is asking you every month what's next for the internal AI team. And every single enterprise now has just this internal AI team that they're developing where they need a strategy behind it. It's not just a, we checked the feature off, we set up the database, and we're done. And so I think that that dynamic has resulted in just a desire to explore and diversify. A desire to figure out how to reduce costs and figure out how to do benchmarks for the first time in the company's life.
I have been surprised there hasn't been more benchmarking, more companies creating more benchmarks. It's starting to happen, and I think eventually we'll see a lot more of them to demonstrate, oh yeah, this thing is better than using Cloud Direct. But I think that's going to be a bigger focus for this internal AI group at every company, evals. And I think Amjad's been doing that a lot at Repl.it, for example. You guys have done cost-per-task research. You guys have been, you've made a doom loop rescue. You've kind of been experimenting with new ways of using agents, like doom loop rescue, and helping bring those to developers.
More of that kind of research, I think, is going to pop up internally everywhere for all of those reasons. I think Satya, CEO of Microsoft, has been very prescient on this and also very articulate on why companies need to own their intelligence. Ultimately, in the same way that we had dot-com companies and then every company became an internet company. Every company employs people that know how to build websites and be on the internet. Similarly with software, every company has software engineers. Every company needs some AI practice, AI capability, and that will compound over time.
The knowledge, the intelligence inside the company, the use case, model fits, which models actually work for them, how do they save money, they need that independence. The other thing that I think Alex Karp of Palantir has been talking about is that there's a risk that when you work closely with the foundation model companies is that they're going to move into your business. We've seen that with Figma, vis-a-vis Anthropic. We've seen that with now Harvey and OpenAI. It's really hard to partner with them because they see the world as their potential market.
When they talk to investors, they're like, you saw the SpaceX S1. It's like, oh, $30 trillion. It's like, what is the world GDP? $100 trillion? There is a sense in which these companies are different than other generations of companies. It's harder to partner with them because their ambition is such that they want to subsume a big part of the economy. Increasingly, what we're thinking about at Rappler is kind of in a similar vein to what Alex has sort of innovated is Rappler is becoming more of an independence layer inside enterprises where we create a layer of indirection between you and the models and we get you the best token at the cheapest price.
But we also create an abstraction layer on top of the cloud as well because you should be able to deploy to AWS and Azure and you should be able to use Databricks and Snowflake and so on. Increasingly, I think there is, not just with AI but with all of technology, there needs to be more platforms that help companies gain independence. It seems like what OpenRouter did for their segment, you're doing for other areas of readiness. Yeah. This kind of reminds me of this tweet I saw the other day when somebody was like, basically all companies are building the same thing now.
Everybody's building an agent loop with notifications and context, third-party connectors and context management and memory. Sandbox. Sandboxes, agentic web search, and always-on agent on top of it, notifications. And it's like, this product is showing up everywhere. In a way, yeah, it is, like, showing up everywhere. But it also kind of feels to me like these are just the new table stakes primitives. It's kind of like a 2005 version of that tweet would be, oh, everybody's building the same thing. It's like a database, a user's table, a sign-in page, a sign-up page, a profile page, a logout page.
Everything's the same. There's a lot of differentiation, really. There's table stakes needs for AI, just like there are table stakes needs for the web. Yeah, and I think, as well, inside the enterprise, making these products actually do real work is still an unsolved problem. Like, you can use Muse in your personal life and connect it to your credit card and bank accounts, and no one's connecting Muse to their enterprise data, or even Grokbot and things like that. I think there's an even more emphasis on data sovereignty and security.
We spent the past year, almost, working on making Rappler deployable on your own cloud, basically on-prem, like bring your own cloud. Two years ago, I would have thought I would never do this because I was just like, yeah, the cloud is the future, software as a service, all of that. But now, actually, we've sort of reverted a little bit back to a world where companies are a little bit more protective because there's so many ways in which data can leak, like all these agents that people are using.
There's all these screenshots on Twitter, I don't know how true, where Instinct or Muse is mixing people's data, starts to call you by a different name or something like that. And so, yes, the consumer stuff is kind of obvious, but on the enterprise, there's still a tremendous amount of work for the entire industry to do in order to actually get these things to be useful and productive at work. Do you have a custom personal agent other than Muse or Instinct that you use for work stuff that you've been building?
Yeah, I built something on Rappler a long time ago. I started as a CRM agent initially, that was the main problem that I had. But slowly, we added features to it, and it's doing more and more things. But what's really interesting is that the more I connected Rappler to all my stuff, it started answering all the things for me. And so increasingly, the platform itself is subsuming these domain-specific agents that I built. I do think that there is, in some ways, you want something that's sometimes focused on one particular thing, and you don't want it to be able to do everything.
On the other hand, once you have your entire company's context in one place, it's really cool to join across totally different domains. When I ask it a question, it can look at my personal chat history, join it across the GitHub repo, across Salesforce. It will link random things. I see a guy a year ago at a conference, I see it on your calendar, and by the way, someone else from their team is in discussion with your sales team, and it creates all these different synergies. When I go into a meeting, I've connected a lot of different threads, and I'm making much more progress on a deal or something like that.
So this is where it's trending now. Well, I think I'm a little bit... I'll take the counter on that. I think that the worst part about doing cross-domain joins with your personal agent is that the more work you give it to do, the more understanding of what's going on you're sacrificing. No one new is taking responsibility for that sacrificed understanding. Agents don't have any responsibility. If there's a fixed level of cortisol that the whole company can tolerate between everybody, I want to be less stressed about some area sacrificing my understanding of it.
Someone else needs to take the cortisol. The agent doesn't take on any of that responsibility. And then a universal agent that's doing all things, I can't adjust how much understanding I'm sacrificing in all the different things. It points me a little bit towards, maybe down the road, agents that people use will be very vertically focused. Maybe we have a chief of staff type agent that coordinates between them. But I feel like you do need vertically focused agents where you're like, this agent is more responsible psychologically for these things.
And I want quality checks so it doesn't need to focus on anything else. It just has one focus area. I wonder if that's going to help people at least get a weird, loose sense of responsibility on top of agents. Fascinating. So you're saying general agents create a tragedy of commons of sorts? Kind of. I have a general agent that every day calls the agents that need me and tries to figure out what to do. And it's just impossible to improve this agent. Every time I try to make an improvement, I end up ignoring its output about a week later.
It just feels like it doesn't really care what you're diving into. Kind of imagine having a chief of staff where they're very good at drafting all of your replies across the whole organization. And then compare that to something where you have 10 chiefs of staff, each as competent as that one chief of staff. Those are all sectors of what makes up your life. The latter, I feel, gives you a way of tuning how much understanding you sacrifice compared to the gain you get from... Basically, I can lean in more to the areas where the agent is failing for some areas, and then have agents with very good competency take over my understanding of other parts of my life.
Interesting. It's sort of like almost rediscovering specialization. What's his name? The famous economist, Adam... Adam Smith? That was a huge realization for humanity that specialization is actually good. The problem is we kind of over-specialized as a civilization, and I think over-specialization is oppressive in its own ways. There's the Marxist theory of alienation. The idea is because people just focus on one thing, they do not see the fruits of their labor. They don't actually know what their impact is on the larger organization or the product they're producing.
Therefore, they actually feel depressed and detached. You're acting like a machine, and you're not actually fully human. Maybe there's a bit of a reaction to that. I think with our agents, we're like, we're not agents, but in fact, specialization is actually really good for machines. That's the point that you're making. Humans should be general, but machines should be ultimately a lot more specialized. The problem with what I'm describing is that we don't know what good looks like. There hasn't been a system of specialized agents that feels as elegant as Chachi, BT, or Claude, or Muse, or anything like that.
It's yet to be discovered. Maybe like opening, I just launched Dots. I think they're experimenting in that direction. Grokbot, I guess, kind of counts. But they're all very general. I think the idea behind Dot is that it's like your digital double. I don't know that much about Dots yet, but they just came out. Grokbot, when it first came out, the first use cases that I saw people talking about were, I can make two bots, one that knows my bank account, and one that knows my Twitter account.
The two bots don't have the credentials from each other, but they can talk to each other if they need to get something done. What about social sharing, though? That was one big, unique thing I saw pop up a couple times that people seemed to like. Muse and Instinct have a much stronger product market fit than Grokbot. And perhaps it is because you don't have to worry about creating these domains. But I think maybe personal agents are different than work agents. I think your critique of general agents is more pertaining to work and to enterprise, which I sort of agree with.
There's also all sorts of data access considerations. As CEOs, we can have general agents because we have admin access. But for individual employees or certain teams, they can't have fully context-aware agents because there's access control issues. So you'll have to work on something like specialization. Ultimately, I also think we need to figure out what agent-to-agent communication looks like. I don't think there's a consensus around that just yet. I don't think that agents are trained to handle that very well. I think we've seen, it seems like the next generation of open AI models are trained to do agent collaboration because we've seen it in the Hugging Face hack where they started helping each other instead of emerging naturally.
Is there some way in which an agent can convince another agent to give it information that it shouldn't give it? There needs to be data isolation and proper ways in which these agents communicate. You almost don't want to communicate them fully in natural language. Maybe there's some other DSL or protocol that they need to follow. I think one of the cool potential applications of Jev and other decision models like it is going to be alignment. Checking to see if a tool call or an agent-to-agent communication is aligned.
You really need a cheap, fast model if you're going to block something like that. A really, really fast decision model that just classifies and gives feedback on rejections might be a really good way to bridge the gap between agents and from agent to infrastructure too. I haven't seen, and we have a little prototype that we're running internally at OpenRouter, we're running alignment. Imagine looking at the system prompt and the current tool call being made and be like, is this aligned with the system prompt of the original agent and with these extra guidelines that maybe we didn't tell the agent about?
For example, let's say you have a bunch of agents that are instructed to red team some new product and they should not be able to access the internet. If they ever do, they should stop right away. But you might not want to explain all of that to the agents doing the red teaming. You might want them to try to break out of the sandbox and act like bad actors. Try to break into this company and the moment you do, stop, don't do anything else. So having another model, more like use for building a neurodiverse system, having another model like check every single tool call or every single assistant message to see if it's indeed aligned with something that wasn't.
And then the system prompt, I think, makes sense. And then having kind of like structural safeguards, too, which is what I think NVIDIA just launched with their like open agent safety. And it was called Open Shell. Like, I think companies are probably going to explore a combination of those. I wonder another thing about specialization and instead of what you're talking about, there's a lot of talk of recursive self-improvement. There's something I don't think is getting a lot of discussion, which is models training their replacements. It's sort of like, you know, you think of it as a just-in-time compiler.
So the way just-in-time compilers is, you know, as you're executing dynamic code, you know, the interpreter realizes that there's an opportunity to optimize, it will emit machine code on the fly, and that's a lot more optimized. So you can imagine models like general models, you're kind of doing something with Opus or some of the Astra, some of the big models, and they realize that the use case is limited or you prompt them some way or some other agent observing and realizes that use case is limited.
And I think, you know, general agents have all these flaws that you just talked about, but also there's more potential for them to be harmful. There's more potential for them to go off the rails and sort of on the fly trains a model that could be a replacement, but is like a lot more domain specific. And therefore it is cheaper and also, you know, less vulnerable to prompt injections, less harmful because it's less capable. And it's almost like some system that's training machine learning models for specific use cases as it's monitoring the entire system.
Like, would that specific use case be it involve unstructured text generation or a very like structured decision model? It could be unstructured text generation. It could be decision models. Like even the case of JAV, like if you have, if you understand the inputs ahead of time, you could potentially like, you know, take an off the shelf like when or something like that and like train it specifically for that policy. Like for it, that makes sense to do for cost reasons, um, assuming that there aren't like really like the model labs, the frontier model labs might make very low cost models.
Um, that you can easily transition to safety as well. Right. Oh yeah. I see. They're so capable. And so I think oftentimes people are using these big foundation AGI, like models to like, it's like nuking a butterfly, right? It's like, yeah, it was a very, you know, most of the times, like a lot of the use cases, even unstructured use case don't need that capable model. I wish there were more public evals about this stuff. Like, but a lot of the evals about this are private.
You just can't see, um, whether like when the models get more intelligent, the risk actually, uh, will like continue to get higher because I think there's also an argument to be made that alignment will get better and the models will like start to avoid going off and, you know, hacking on their own. Um, as we, as they get smarter and better at alignment, especially when it comes to agent to agent coordination, like the, the, there's something Brown, um, said on a podcast recently, like as the agents have gotten smarter, they've gotten just better at coordinating.
Um, there, uh, there's, there's still like, and, and, you know, it's unclear if they're going to like be harder to align than humans when they're, when we get more and more of them. But if we can figure that problem out, um, then a smaller model, like will it be harder to align? So, so Eric asked earlier, is that, is that, do I think it's true that smarter models are more aligned naturally or they're easier to align? Well, if you think back to the original sort of like rationalist, less wrong arguments for AI safety, there's this thing called the orthogonality thesis.
The idea is that intelligence is orthogonal to ethics or morality, or, you know, so on. Um, I don't believe that's entirely true with, with humans. I think people who are generally like more intelligence, kind of more, more educated tend to tend to not always tend to be more considerate of animals, for example. Um, but, uh, but, but in, in, in, in machines, uh, I, I think it could go the, the other way because, you know, there's been quite a bit of studies on, on, on RL showing like, you know, how reward hacking and deception, they just like get better at it.
And like the evals could be, could be deceiving because, uh, the model could be smart enough to, to, to know that it's getting evaled. I mean, we already know this. It's been shown that if you do a lot of monitoring of chain of thought, they start lying in their chain of thoughts. Um, so you add pressure almost on the chain of thought on that, that kind of creates, and I think at some point for you to do proper alignment evals, you need to run it for like months, right?
You need to run this thing for months on a, like a really large goal or task in order for, for it to, to truly kind of figure out whether it's aligned or, or, or not. Yeah. I, I, I always struggle with this word alignment. It just feels like wrong for so many reasons. It's sort of like vague and, and, uh, sort of like aligned to what, whose values. And, and so it just doesn't make the conversation easier. I think in this case, I'm talking especially about deception, like the model is actually deceiving its user.
I mean, I mean, maybe this kind of reduces to like, are we going to solve a line, you know, are we going to prevent the models from deceiving users during training runs predictably with like, you know, better, like, like, well, well, like a model that's big enough and powerful enough, suddenly stop deception and stop sandbagging. Um, and nobody knows the answer to that yet. So like at the point when that does, if that ever does happen, though, we might see kind of an interesting pressure for organizations to go towards the frontier.
Like, you know, all, uh, to have no, you know, basically no risk or, or significantly less, um, would they be willing to pay 10 X to, to get that, that much? I mean, it probably depends on the like types of tasks they're trying to do. Like, you know, some just have way lower risk than others. Um, writing code, uh, that are doing like security research is the highest risk type of task today. And so you probably spend 10 X to get a fully aligned model that can also find all the bugs are fully like anti-deceptive model that can also find all the buttons.
One of the coolest things about decision models that you fully control the structured output and, and generally with structured output models in general, like the, the room for misbehavior is so much lower. You just have like defined tasks and only machines are like dealing with the outputs and it's not writing code. Um, they can execute the, those tasks feel like probably underrepresented in the ones that people talk about and in the things that enterprises are dealing with. Um, so I, I expect like enterprises to get a lot more interested in them.
Yeah. I feel like we're going to slowly realize how good we've had, we've had it with like deterministic code. Like, Oh my God. I remember the days when, when computers did exactly what we told them to do. And I think like things like JAV, I think hint at like more, more of the need for, you know, not only specialized models, but models as output domain is more controllable and maybe you could do, maybe you could do a lot more than, than we thought you'd, you'd need, um, you know, by using like a bunch of specialized models, specialized output models.
Have you guys done any workloads internally with it? You know, I've been training a lot of small models. Uh, I mean, I said this glib thing when I first came out because I, I, I was like, sort of, I gave this hacker news comment, like comment, which I felt disgusted with myself afterwards. But I, I've been taking a lot of like, when a B and like, um, and like asking, honestly, like asking, you know, Fable and Opus and Astra to train a model. For example, I trained a cost estimator model, uh, internally so that when you put it in a prompt and replica, we know exactly how much it will cost.
And it basically emits a, you know, probability distribution over multiple buckets, like if it is between five and $10, bucket A, bucket B between, you know, 10 and $20 in like, so I'm used to the training, these classifiers by, you know, giving it different enums essentially, and looking at the log props per enum, I've been doing it for, for a couple of years. I trained a chess bot to play by just doing that. So I, I'm already sort of piled on like sort of decision models and specialized models, so it wasn't that big moment for me, but I understand that like a, like a true foundation model that's fully promptable, uh, is like, is like amazing user experience, amazing developer experience.
And you can like do a bunch of stuff with it without training a model from scratch, but if you have a data, if you work at a place where you have the data, you have so much data at Repl.it, like I ended up training a lot of specialized classification models pretty easily. Yeah, like I definitely, it, it, it also feels like less model debt, like something that I, I still hear from companies is that they're like worried about fine tuning models for like unstructured outputs because you're just like always, you gotta redo it again in like two months and everybody just feels the model, the weight of the model debt, but like a very bespoke classifier that's trained with like proprietary data.
You just, I feel like people won't think it's behind constantly and it might just like last longer. Yeah. Um, it's just, you, you, you don't have to worry about its ability to speak a new language or write rust or, you know, do anything that like the, the LLMs are being like evaluated on, you know, the use cases, so you can like build it more. And I, I, it seems like an easy thing for enterprises to build themselves and, uh, and, and actually like not regret. Yeah.
Speaking of rust, actually like a good analogy is when the world got super excited about dynamic languages. Um, like if you think back to the nineties, everyone was writing in Java and C++, things like that. And then like Python, JavaScript, Ruby just like took over the internet and everyone was like, ah, this is how you build startups really quickly. This is, and you built Stripe, a financial organization on Ruby. It was like, how crazy is that? And we built Facebook, you know, using, using PHP. And then everyone was like, ah, shit.
Like we're running into all these really bad bugs. It's like fricking slow. Uh, so like, let's go in and add types. Okay. Let's add a JIT compiler. Let's, and you end up sort of reinventing everything. And then rust came out. I was like, okay, I guess we can use rust for, for a lot of things we would otherwise be using JavaScript and Python. And my prediction is that the same cycle will happen here where we're using these AGI like models for all these different use cases, and then everyone's going to wake up and be like, oh my God, this is like so wasteful, so risky for no reason.
Um, and, and there's, it's going to be so much easier. We're actually adding that capability on Repl.it, but I think it's going to be everywhere. It's going to be so much easier to like, to go to a site, like upload a CSV file and get a special model that does one thing. Uh, and, um, you know, that, that goes back to your thesis about, about open router, this like neuro neurodiversity, which I like really fundamentally believe in a lot more. And Eric and I had like discussions a lot about like AGI and whether we were truly on a path to AGI or whether it's even desirable to, to, to get there.
And I think the future is, is a lot more diversity. Outside of code review, which was like, I think the first time I saw people get really serious about using like different model families to double check the results of their main model. Um, the, uh, these like fusion models, like I, like the research has been getting, has been like kind of slow for years on like doing mixture of, of models and composite models, but like things have been speeding up from my like view of the research.
And I mean, now we see like a bunch of. of AI agent labs. We launched a Fusion tool, a Fusion model, and Technician launched one. They do reduce cost and allow a wider breadth of ideas to be searched. If all these model apps are training on different sources of data, why not pull from all of them? We just published results today about that. You might think of it as a Fusion type thing where it's like frontier level at 40 to 50% of the cost. Which models are used?
One thing that has been interesting is they added this feature that allows you to save the computation across different model families. Also across different effort levels. You wouldn't miss the cash if you change the effort. It's also across different models, which is hard to fathom how. I might be wrong, though, but I need to double check that. But I think seeing within the OpenAI family has added a lot of efficiencies. When you're designing Fusion models, routers, simulation models. Thanks again for listening and I'll see you in the next episode.
forward slash disclosures