1% better

Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper

Treat AI agents less like chatbots and more like teammates: give one a clear, high-level task before you finish work today, let it research or draft in the background, and review the result tomorrow. This shifts you from a slow ask-wait-reply loop to asynchronous delegation. Start with work that ben

1h 23m

Summary published by , updated .

Invest Like The Best

Key Takeaway

Treat AI agents less like chatbots and more like teammates: give one a clear, high-level task before you finish work today, let it research or draft in the background, and review the result tomorrow. This shifts you from a slow ask-wait-reply loop to asynchronous delegation. Start with work that benefits from breadth rather than instant answers—competitive research, a source-backed brief, documentation review, or a list of potential risks.

Episode Overview

The guest describes Sal Research’s goal of making open-source AI inference dramatically cheaper by optimizing every layer of the stack: software, chips, data centers, and energy. The conversation argues that the most important AI shift will be from low-latency chatbots to long-running background agents, while exploring the technical and economic infrastructure required to support that shift.

Key Insights

Background agents beat constant interaction

The guest expects AI use to move from real-time chat toward agents that work for hours or longer while people are away. In this model, humans set high-level goals and review outcomes periodically, rather than becoming a bottleneck by responding every few minutes.

Cheaper intelligence enables new behavior

A large price reduction is not merely an incremental improvement; it creates a new product category. When tokens become inexpensive enough to spend without a guaranteed return, proactive assistants can continuously monitor information, research broadly, and surface useful next steps.

Optimize for throughput when the task can wait

Low latency and high throughput require different system trade-offs. For background work, batching more tasks onto hardware can lower cost and increase utilization even if an individual request takes longer, making slower but cheaper infrastructure economically attractive.

Verified environments can drive self-improvement

The guest argues that models improve most effectively on tasks where progress can be measured automatically, such as coding and mathematics. Give an agent a defined environment, a measurable objective, and feedback on whether it succeeded; the environment itself becomes useful training data.

Unused compute is an orchestration problem

The guest sees idle chips, smaller data centers, lower-reliability capacity, and intermittent renewable energy as underused supply. Rather than relying only on the newest hardware, the opportunity is to match each type of compute to workloads that can tolerate its constraints.

Frameworks or Models

Speed of Light

1. Identify the hardware’s theoretical performance limit. 2. Measure every bottleneck that prevents reaching it, including software, power, thermal limits, and communication overhead. 3. Optimize against the absolute achievable limit rather than against competitors’ relative performance. 4. Repeat as new bottlenecks emerge.

Garbage Collection Strategy

1. Find underused or underpriced compute capacity, including less-popular chips and small data centers. 2. Match it to asynchronous workloads that can tolerate interruptions and variable latency. 3. Use orchestration software to move work when a site or machine fails. 4. Aggregate this dispersed capacity into a lower-cost inference network.

Verified-Task Self-Improvement

1. Place an agent in an environment with a concrete task. 2. Define an automatic or objective way to measure progress and success. 3. Let the agent attempt, receive feedback, and iterate. 4. Expand this approach across domains with verifiable outcomes, such as code, mathematics, and other formal problems.

Notable Quotes

"I want it to be proactively. I want, so that it is in background mode. It can be said that best delay– this is her full absence. When you you wake up in the morning, the work is already done overnight."

— Guest

"You are not manage your own colleagues every 5 minutes. You ask them to complete the task high level, and then you come back and checking, maybe every day, but rather once a week."

— Guest

"You must be willing to spend tokens without any promises of return. This is the key to success."

— Guest

"There are no bad chips. There are really bad prices."

— Guest

"We still treat agent as a person, consultation with which are expensive, and you must ask him when you have a difficult question. It wrong way of thinking about intelligence."

— Guest

Action Items

  • 1
    Delegate one overnight task

    Before ending your workday, assign an AI agent one bounded, high-value task: specify the desired output, constraints, sources or data it may use, and a definition of done. Review the work the next morning instead of monitoring every step.

  • 2
    Separate instant work from background work

    List your recurring AI use cases and label each as either real-time or asynchronous. Move research, document analysis, idea generation, testing, and monitoring into the asynchronous category whenever a delay will not reduce the value of the result.

  • 3
    Use verifiable goals

    For agent tasks, define an objective that can be checked: tests passing, required sources cited, fields completed, inconsistencies flagged, or a rubric score reached. Clear feedback loops reduce vague prompting and make iteration easier.

  • 4
    Build an intelligence budget

    Choose one project where deeper analysis could create outsized value and allocate a fixed AI budget for it this week. Judge the outcome by the quality of the decision or deliverable—not by how few tokens or prompts were used.

Full Transcript

Transcript of Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper from Invest Like The Best. Auto-generated from episode audio; may contain minor errors.

My job is to in order to do tokens as much as possible cheaper. I will achieve this, and I will do it through each level in stack available to me . I love levers. suggestions. I will use each chip. I will use every source energy and each a piece of land in United States, which, you know, suitable for this. We still treat to the agent as to person, consultation with which they stand expensive, and you should ask them when You have a difficult question. . This is not the way.

thinking about intelligence. Unbelievable, what the machine can do to think, and we must to try to convey this to the greatest extent possible number of people. I I think that at the beginning these conversations are important just to say literally, what are you create and what it is is doing today. So, maybe just guide us there by means of short description, literally what is this the system you create, and why she must exist. Sal research is a factory tokens. We have an API where anyone can send ask us where they are can use large language models, large language models with open source for any task, what they want.

Ahem, we will provide them with these tokens at a price that has no equal in market. We also we support their ability to create agents based on this. We place what we call sandboxes, that is virtual machines agents of long-term works placed in cloud and dedicated for agents who working for hours, days or weeks. Therefore you must to think of oneself as about a peer company for others who serve various types of logic conclusion. You serve one specific type logical conclusion, and your goal is to be the cheapest supplier and those who provide a certain kind using intelligence.

That's right. Our topic companies—prosperity. We want to deliver this new product— intelligence—as much as possible more people at a price that is stable for almost each industry. We we believe that when you do something 10 times cheaper, it's new product category, and we strive to do this is for tokens. We we believe that this so deep that a machine can think and now our task —to force as much as possible more cars in the world to work on thinking. So, if you think, uh, about topic of the day–value tokens, is it worth it?

start with the cost tokens? Is there any another way put it on foreground? Absolutely. Of token value today, it's mine a guiding star is to have the lowest cost per token in industry, and to do this much easier. I I don't think tokens are– this is the final unit, e- uh, work or intelligence, but precisely we use them today, so this very simple. I think that after tokens you start moving more to, uh, more quantity results that are in a vague direction. Can you imagine? for example, today, when you consume tokens through an agent, you really don't control how much token agent justifies.

He can justify this during a certain period time or maybe to cause a certain number tools; and that's it more often, I think, in we will be agents, which will perform a certain unit of work, will do their best more blows to gate, and how many they are not tokens used, this will be of a kind dependent variable, which depends on task. So, you thinking about agents, which independently manage the budget tokens, unlike from a company that sets a budget for how many tokens engineers can spend every month.

Why is there a possibility, which you can use? It seems that the whole the world now oriented towards a larger number, better, faster and cheaper tokens. It seems that the world trying to decide this problem very aggressively. What a unique opportunity you saw, perhaps, because the market ineffective trying to decide it? So I think there is two things that are a tailwind for our company. One of them is growth popularity open source. I should have talked about it first. I think that we begin to see, as more and more of our customers and wider the market cares about possession of intelligence .

They want to have sovereignty of control over what depend. Uh, and this created much a stronger market for customized models or even just like these vanilla models with open source, which no one will ever be able to take it away from you. You have there is always weight. You you always have the right deploy them as you wish I like it. Uh, in in this world during the last few existed for quite a few years a strong market that served these models in large scale. Problem is that all these companies, you you can choose, decimal, fireworks together, all of them focused on low-latency logical conclusion, and dragged them into this one direction very much important client, uh, cursor.

And I think that that was correct choice about a year six months ago so it started to seem that, maybe low Delay is not the only thing, What do you want from agent. You wanted greater stability, more long-term tasks. And now for It's quite obvious to me, what is the future agent logical conclusion is long-term task. You are you going to launch car for hours or for days in a row. It doesn't matter whether it gives out she tokens with at a speed of 100 tokens in a second.

Maybe 10 quite enough. AND this is accompanied by appropriate advantages and efficiency. Why are you so sure? in this? Me It seems that I want to so that everything is as good as possible rather. When you are waiting for this, you absolutely deserve as soon as possible answers. My trick in because I don't want to You've been waiting for this. I I want it to be proactively. I want, so that it is in background mode. It can be said that best delay– this is her full absence.

When you you wake up in the morning, the work is already done overnight. You don't even I needed to talk about this. ask. Eh, this is a dream. We're not quite there yet. But more importantly, I I think that the more you are in cycle when you tell agents and waiting for an answer, the more you become a bottleneck, helping the agent to perform more- less work. We would like an agent worked in more than human time scale. You are not manage your own colleagues every 5 minutes.

You ask them to complete the task high level, and then you come back and checking, maybe every day, but rather once a week. And this the future for me cooperation between people and agents, rather human skills time. Tell me more about early signs that it is is happening, and therefore you should build this company. So, first and the most important thing is idea of ​​scaling computing time testing. Uh, idea. about what can be given the agent has more time, and he will give you the best respond.

It was theoretically discussed about 2 years ago, but actually we couldn't to rely on this, so far, I would say, not yet Opus 45 appeared at the end of the past year. Opus45 was the first agent who is generally suitable for the tasks with long-term horizon, and, you know, he was quite mediocre when first appeared. But if you look at newer models, and also that we done with open code, it is clear that agents are capable to work during hours. I wouldn't say that these are days, but an hour exactly, completely suitable today.

So, just seeing that middle queue or task duration are becoming more and more longer, no need many points to to display the exponent and see what agents are worth it to launch them for longer periods periods. Which one, on in your opinion, will be agent market share with a long period in 3 years or something similar? You know, I love this one. market because it unlimited. In a cycle there is no person. So, you can you use as much as you want tokens in the background mode.

Uh, in compared to duration of human attention. If you tell me use in 10 times more tokens on CodeEx or in Cloud Code, I not really I'm sure I can. yet. I'm already hooked and locked in coding more part of the day while I'm away laptop. What else unknown, so that's it, how many tokens can use in in the background or proactively. So in long-term in the long run, I think, you know, we'll finish this year, perhaps with working 50/50 load in in the background and in real-time time, but I see that it will increase to 90/10 on background benefit regime.

What are your favorites? examples of something that much better is performed in background than in the task of man, Which one is writing? Most deep research, most of the questions where you want to get the final answer not on 100 sources, not on a thousand sources, and 10 000 sources or more. If you want create authoritative index information, such as, for example, one of our customers, who is engaged in parallel web- systems, strives create index by the entire Internet and monitor changes in Internet in mode real time.

It such a crazy task in exabyte scale, for whose achievement absolutely necessary different type or scale intelligence. Deep research is ours main category, and we are increasingly seeing like cybersecurity moving in this direction. If think, yes, you can generate like this a lot of code, but exists exponentially more ways break the same code than generate this the code itself. And there are wonderful ones customers who are very working hard to find agents, which can break any software provision and proactively it to fix. So, when, for example, for the first time Fable or Mythos came out, in community cybersecurity was desire to launch Fable on every line the code we used to wrote something, and look for errors 20 in different ways.

It means that you looking for errors memory and errors business logic and network vulnerabilities , all of this. And that's all, why would you wrote specialized agents. You wouldn't just once wanted Fable reviewed the weekend code, you would actually created environments, where can you check these programs on the presence of a pentest. AND at some point people started joking that security has become proof of work. When you need secure software software, actually the question is because how many dollars you spent on anthropogenic APIs, trying to break your software software.

Uh, this is best indicator how much it is safe because it is the best tool in the world. And we all more often found that the limit of intelligence is here quite narrow. This is not the case when Fable finds a superset all errors in software provision. You would found some errors with very little model that you don't you will find with great model. Uh, you would found some Haiku mistakes you would found from Fable, and vice versa. So this is encouraged this very diverse approach to sample and try create agents cybersecurity, which autonomously break software ensuring that you could fix them.

If you started to speculate and imagine what things can implement very cheap agents for long-term using. We talked about some very practical examples of deep research, hmm, cybersecurity, etc. But if you are a little dream about options using the new product categories, which this view will reveal conclusions, and I think The question is: "What, if you had reached maximum success , dream a little about what it can realize?" Yes, of course. So I I think that for some users most of all me the idea is fascinating proactive intellectual agents.

You can imagine Siri, who constantly working in in the background to understand which emails you received in a day, all textual messages you received in a day, and has much more encyclopedic a look at your life and how to be useful in this life. Now it's still there point solutions, therefore you have to to do a lot tips. Siri not so much proactive. This is what we can fix it with the help of abundant abundant conclusion. If you trust the car so much that she is reliable and trustworthy, as in private life , you can even imagine that the car can understand how you interact with it, and proactively detect your next steps.

Whenever you open the phone, can we to build a good one model of what you are you going to do further? I appreciate that yes, we We can. And the key to this is incredibly cheap intelligence. You must be willing to spend tokens without any promises of return. This is the key to success. A distant view of this is that in we are a form of intelligence , which can solve any verified problem. Any proven problem means the majority software software. It means a lot formal, such as mathematical proofs etc.

And this can also to mean scientific discovery. And all this relatively proven problems. And all these things currently have its price in dollars. This is hidden value. How many tokens you could use it to Did it work? And we actually started to consider it as reasonable price for these long-term tasks. This is not millions, that's thousands, and perhaps the nearest sometimes it can be hundreds or even tens of dollars to to get the final answer to any scientific question or research problem. RAMP is a single platform, created for to make your financial team more flexible, faster and better, saving business on average 5% per year, so that you can to focus on growth.

RAM clients increase income by 3.2 times faster than average American business . Visa, Verscent, Kerser, Stripe, Notion, 11 Lab, Shopify and 70,000 other companies now working on RAMP. My the company also works, and yours too must. Find out more at ramp.com/invest. OpenAI, Cursor, Anthropic, Perplexity and Verscent have something common. All of them use working OS. That achieve a large-scale implementation on the company, you need to provide main features, such as SSO, skim, arbback and audit logs. Instead of spend months on independent creation these critical important opportunities , you can simply use API working OS, so get them all from zero day.

Here why so many leading teams from artificial intelligence, that you hear about, already running on Work OS. Work OS –this is the fastest way to prepare to work in enterprise and to focus on, what is most important– your product. Visit works.com to to start. Felix by Rogo– this is a personal agent finance, which transforms one request for ready-made work ready for client, using own templates, context and your standards companies. Send Felix the electronic a letter like: " Take these comments and "Process them for me." .

Or "Update my tracker by context these electronic letters. And Felix will send back ready PowerPoint presentations, Excel models and research from sources. Felix works like this yours already does this team, performing work quickly and exactly around the clock. Learn more at rogo.ai/felixelix. And therefore, if we dream about such a future, we then we limit ourselves only questions that can put people. Essentially, questions that we we can put Um, models on the verge of to take even question of high level and to pursue him with the help of everyone possible further actions, you can ask the model for to actually take on this independently, and the question is the one you have token budget, and we will decide budget problem tokens, and how about unverified tasks?

I mainly refer to the entire category human taste for of this category. We still did not solve the problem human taste, and I I don't know if this is it. It is possible in principle. I am interested in being surprised, but, hmm, we are focused on very quantitative, hmm, problems; we we leave quality letters, we, hmm, beauty art, people. Good. Now let's let's talk about, um, very smart set decisions you make hope create in the end, to have this giant factory tokens, extremely cheap supplier extremely cheap intelligence.

I think you think about this from the point of view software level software, hardware provision, hmm, and power. Tell me what yours is. master plan to approach this a challenge that is so differs from what others think. You know, we always should start with software software. You know, where there is an opportunity in modern chips with modern centers data processing to increase efficiency? And the first what we did is tried build the whole LM software stack around the peak efficiency graphics processor . That is, we we use graphics processors Nvidia; we wanted squeeze more tokens from one chip than anyone else in world, and this starts with lowest level software kernels.

It actually my experience; I have spent my whole life actually, all my own professional life, working on graphic processors and nuclei. Nvidia had mine first job when I studied in college, and, uh, I saw how tensor kernels have earned their right to be on the chip It was back in 2016. Just describe what it is. means for the average person. So, okay, tensor the core is specialized unit on the graphic processor, which accelerates multiplication matrices. That's it. simply. There is a long the story of how we developed in tensor kernel with the time we Let's consider.

And why multiplication? Are matrices so important? That's a great question. Actually I can't. to claim that there is some divine truth inverse, which explains why matrix multiplication seems atomic unit of calculation. But, uh, one of the the ways in which I do this described, this is what it is very concise a way to combine two blocks of numbers together and to force them interact somehow in an interesting way. That's all, What can I do about it? say. Very convenient, that linear algebra turns out to be very compact presentation arbitrary relationships in data.

So wonderful graphic company Nvidia, obviously, dominated the market graphics processors and game graphics for quite a while for a long time. AND then, starting approximately from the middle 2010s, they started run these strange projects to do graphics processor more suitable for machine tasks training, which they tracked. I I remember reading. some laboratory some of my notebooks managers when I worked at Nvidia. They visited these small conferences from machine learning, such as ICML or NURPS in that time, and just took these into account article, saying, "Oh, this a thing with a deep learning, it seems, is gaining popularity ".

And what's really interesting , so it is that these postgraduate students use gaming graphics Nvidia processors for teaching their own large models. To us should be double-clicked on this and find out, What is happening here? And by 2015-2016, at least Jensen was convinced that need to double effort: "Hey, this using our models only for our chips will be only grow; let's start to allocate more and more and more precious silicon area crystal for this opportunities, which, seems to appear ". Let's place first version tensor kernels on chippy.

So, we we are talking about take this game chip , which is intended for drawing pixels on the screen, and adapt it for execution metric multiplications, and it was early stage, and you would competed with graphic commands developers, in fact, whenever you asked more silicon area and any manufacturing company chips. There is always a reason for this. competition. This is what designers yes carefully guarded. You never want to invest in wrong technology because these are alternative expenses that you could to allocate for some another functionality.

AND that's why we fought tooth and nail, and received only a tiny piece, maybe 5-10%, something like that like this for first generation these chips to to get certain acceleration for basic convolutions, which were fundamental operation for models computer vision on that time. And then we have there was a team programmers, which tried to squeeze from the chip all possible productivity. And I I think that in this to the team of programmers where I work, this is it taught me the most me about, um, just about the ethos that Nvidia has , they have this a term called " speed of light".

They are always chasing. at the speed of light for any the equipment they produce. It so deep rooted in everyone's consciousness engineer, what if the machine can do it to do, we will to push the car forward limits until it is will do what we we consider it possible. And the speed of light is the limit of the possible. The speed of light is the limit of the possible. Exactly Yes. Uh, if we we think that the chip can to work on this frequency and create so much- then multipliers per cycle, we will achieve this.

We are going overcome each bottleneck and to peak productivity level . And that's why I still say to this day to all its engineers, like we're chasing 100 %speed of light. I don't care. relative numbers compared with competition. Me only care about absolute numbers. What we can do with chip and how do we do it will we achieve? Before we leave this section of your being at NVIDIA, that besides this point cultural contact, really changed your thinking about things or stood out the most in how then business was operating or his culture?

IN I have a lot of stories. about Nvidia. We can... we can tell you have a few of them. Hmm, one of my favorites -this is what from the point of view permanent stay at position, many people, with whom I worked in Nvidia in 2015-2016, still there. This company has an incredible number of dependents , and these are the best engineers, honestly speaking, at least from silicon department, with whom I worked for my entire career. They extremely motivated and fascinated. They believed in parallel calculation as the concept in its various incarnations and loved to watch chip development.

It their life's work, and they are extremely competent in this direction. They also very economical company. Nvidia and all others, I think, all Silicon company valleys after 2008 had some reduction and similar benefits. So none free lunch. Uh, for example, Nvidia took another step. Not in the fridge was free milk. So, if you wanted to drink coffee at Nvidia and wanted milk, you really had to every month donate a dollar to dairy club, and the dairy club held Costco's milk refrigerator. And I I remember it clearly.

We we don't do this on sales, but, uh, it thriftiness, which permeates the company. AND therefore, based on this time, you you get this experience of how it is- develop more effective using basic equipment by means of software software. Yes. And so connect this with, you know, modern environment. Yes, absolutely. So, I I think it's graphic. The processor is, in essence, high-speed car bandwidth . Graphics processor ( GPU) the happiest, when you give it to him a lot of work and allow him to carry it out with peak using their computational blocks.

But in reality this is not the way to go we perceived AI the last few years. We really pushed forward AI as interactive chatbot tool, which is the most common form of use AI today. And in this You care a lot about the world. about how to to issue faster answers to a person for keyboard. Of your point of view about what is not needed force user wait, I I want everything to be fine. as soon as possible. And this actually quite interesting for graphics processor .

Very difficult to put a graphic happy processor the path of the complete using computational resources when you trying to be fast issue tokens. There is fundamental compromise for graphics processor between orientation to bandwidth or optimization delays, and all chose optimization delays because form of use was focused on chat bot. I believe that This is the most profound change, which we will see next year. We we are going to move from chatbots to more proactive or background agents. And in this world much more logical build a stack around the checkpoint abilities.

Can you technically explain why compromise between bandwidth and the delay is unbreakable? Why us we can't have both on one device? This is enough. fundamentally almost in every system that could you ever would consider. Always there is a compromise between as soon as possible to miss a small amount of data through the system and to leave a lot buffer space for this, or an attempt to work wide and slow, how narrow and fast, or wide and slowly is classic compromise in all of computer science. But what about specifically graphic processor, I think that there is one thing on which should concentrate, and exactly the concept batch processing on graphics processor .

We want group work many users per package, which we can run simultaneously on the graphic processor. It parallel processing on the graphic processor. We wanted would have a lot parallel work. The thing is, you you do a lot more work when launch a big a batch of calculations together. And thus, you can fill all units, but with each step by step when you execute the package works through graphics processor, yes more work that need to be done. AND therefore any a separate token or request of any individual user in this package will spend more time for graphics processor, transferring it along with traffic other people.

Perhaps, way of saying it is this : if you want get to the center cities and San Francisco, you can take the bus or use private by transport. AND private transport will have its own straight path in a straight line or, you know, using exactly those roads that you want from point A to point B. Bus should serve much more people, and he must to do in principle something that works for everyone, so he chooses slower way, stops and waits, while other people They sit down and get out.

I I think the analogy bus and the car is enough accurate, and that's a great analogy, so the first step of that what are you trying to do to do, to create the best possible bus on graphic Nvidia processors. It your first step optimization. Exactly Yes. This means that we we investigate the following things like different schemes parallelism. Perhaps, this is another example, what can I give you to bring, this, hmm, from graphic Nvidia processors, one of the things they really implemented and with whom it is great managed, -this NVLink interconnection between graphic processors.

AND in fact, this system NVLink is so good, what can you do if you You are a great multiplication. matrices, which you want to perform rather, actually cut it matrix multiplication in half and divide it for two or more, let's say, to eight graphic Nvidia processors, and to force them all to work on parts of this greater multiplication matrices and connect your results together in ends. Reduce their results together in ends. And this is wonderful. way to shorten minimal delay operations. Everyone graphics processor now performs, say, 1/8 smaller work, and therefore he may end faster, but not at eight times faster.

It sublinear scaling. You will be use in eight times more equipment, but not will receive eightfold speed. You can gain speed at about four- five times. You are not get strong scaling. And this through invoices communication costs. It because everyone graphics processor will be a little less efficient, working on a smaller tile work than on larger tile work. So this is the only one way to speed up work if you want minimal possible delay. You can you do it but that's not the choice, which one would I do?

For example, I would prefer use another parallelism scheme, such as expert parallelism or conveyor parallelism. And we we can do interesting things things to overlap and hide communication delay so that you have will be less opportunities to do this is for a server with low latency. So is it right? think of NVLink as technology that improves productivity delays? Yes. And only productivity delays, which will go into next segment that we, like you, you know, we do to another as a company. But yes, NVLink is there mandatory, I would said, for output with low latency.

So Nvidia is great copes with output with low delay. And I say you, what do we really not really worried output with low delay. So where is it? leaves us? Well, I don't. I think I'm late. breath, hoping, that other companies quickly deal with NVLink. This is a complex technology for development. It's hard. to scale. Its difficult to implement in production. Ago, if I have a chip another supplier , and he is fine copes with basic computational components, it is all one can be very good to perform metric multiplication.

He just can't quickly to pass these on results with your analogues. Well, maybe in mine the stack has room for this other chip as really very good calculation option for a dollar. And precisely for this is me actually optimize in in most cases number of flops in this chip and how much it will cost me his hour operation. Uh, yes. other chips that are accurate have a higher rating, than Nvidia, in terms of quantity flops per dollar, but they may not have so much relationships. Ago my job is to in order to find out, what scheme I will be parallelism to use to make this chip suitable for logical conclusion.

This will not be tensor parallelism. Nvidia in basically mandatory for this . But other methods may be a good fit for me. So, first than we will leave part of history, dedicated to the delay, can you comment on the following companies like Cerebras or others who can perform incredibly fast operations? I am interested. What do you think about these? approaches, these companies, what can happen in future. What is yours? forecast for future equipment, oriented towards very low latency? Cerebrus Grock, er, and more several others who are now leaving stealth, I think, made a very interesting not just betting on creating another one graphics processor , and on the creation of another type accelerator, which focused on another hierarchy memory.

Uh, they want to maximize the amount of SRAM on the chip and use it how very, very fast memory for weights odds and cash KV. So, SRAM vs DRAM is two ways to create memory for the chip. One of them - integrate memory on the very logic crystal. So you're telling TSMC: "I I want so much. megabytes of memory on to your chip". Uh, yes, it is way to do it. TSMC has a standard library of cells, which you can use, and you you can just print a bunch SRAM cells.

Problem with SRAMM is that, that it takes a lot places on silicon crystal. So, if you want to collect large crystal, for example, Nvidia Blackwell, with an area of ​​800 mm². If you create all this DS RAM, it will be, perhaps, a few gigabytes, that's not a crazy amount for data storage. Compare this to the fact that if you are ready use completely different technology. So already not TSMC but Micron SKH Highix Samsung . They create DRAM, which is completely different by construction method memory, which is more focused on capacitors than on transistor elements.

So, SRAM, standard way SRAMM construction is what called 6T transistor cell. It stable transistor circuit, which allows write a beat into it, and then she saves this state in this bit regardless of whether are you continuing to serve it. Well, you had to submit some power, but , uh, she's holding this bit without any active management. It is static. Dynamic operational memory (DRAM)—this dynamic memory, because for the record data you record charge on the capacitor, and as soon as you record this charge in capacitor, charge dissipates, he leaks.

So dynamic part of DRAM is that you it takes approximately every 50 milliseconds update every recorded bit. So you constantly juggling billions of balls in in the air, in fact, has billions of bits control controller memory that reads and updates every bit on DRM. The advantage of this is that you you can get much higher density, and this completely different technological technology. There is many different compromises. That's why we divided production dynamic memory (DM) to a completely different one a company like Micron SKX and Samsung.

These are the best companies in the world for this. They produce DM. And if take DM from these companies and conclude its, uh, a lot layers, and print or solder them around the main logic crystal, which you get from Nvidia, now you can get hundreds gigabyte, uh, for example, Blackwell has 288 GB HPM capacity around the logical crystal. And myself logic crystal, maybe it only has about 500 megabytes SRAM. So, perhaps, difference in density DRAM vs. SRAM is several orders, three orders . Okay, so Let's go back to Cerebras.

What are they doing? Well, they see this no problem actually obvious way to increase SRAM density on the chip. But the thing is, SRAMM physically so close to logic gates, which are actually perform calculations , arithmetic-logical the blocks are located right next to SRAMM, with which they brothers are going to data, computational blocks that perform matrix multiplication, can take data from SRAMM with stunning speed. You know, Susquits has a speed of 1 petabyte per second, and Scale Engine 3—21 petabytes per a second later. So compare this to HBM on Nvidia's black wall, this approximately 10 terabytes in a second in this range.

So again anyway, the difference is many orders, larger capacity, but proportionally smaller capacity. So Cerebrus says that we let's take as much as possible more like this crystals. We are not let's limit ourselves to a limit 800 millimeter mesh, TSMC limit of 800 square millimeters, which we imposes TSMC. We we will take all plate, and each the crystal will connect with each other crystal through marking lines. And we let's just try place as much as possible more SRAM throughout plate. And we can to reach, say, 50 gigabytes of SRAM on plate, and then we we are going to fold many plates together in an assembly line or something similar, and now we we can have, you know, up to a terabyte of memory, very, very fast memory, and you do all this work to get ability to read data from SRAMM with speed, yes, 21 petabyte per second to plate.

So , now you can serve these language models with extremely high token speed per a second, because you you can move the whole amount parameters of large models like Kimmy. THERE ARE -uh, you can move all these data to and from the chip, or, sorry, in logical cores and from them approximately a millisecond or something like this. So, you have a way to thousands of tokens for a second. And what is your prediction? for this segment market? Good. So, I think, what about them something is happening hybrid result, for example, us had to combine the Cerebras chip, which is very strong, very good provides fast memory access, with something that has more memory capacity.

Because this it's true that you can take a model with one trillion parameters, such as Kimmy, and place her in large quantities Cerebrus plates, but not you can easily do something do with KB cache. Cache KB grows when people more use model, and it always dynamically. You even you don't know how much KV cache you will be needed. It depends on What users do you have? how much do you have users and how many users you want serve. Or can you explain KV cache is simple, as in basic form?

Yes. So, KB cache when you use language model, each token, which one are you sending through the language model, actually remains in context window language model to those as long as you lead conversation. So, if to talk about 100,000 tokens, then the 100,000th and the first token is still is in conversation, uh, behind us, and the model refers to the whole history of the past conversations to better to predict what will happen further. So, this cache KV– that's a lot of memory. Ahem, you need save presentation for each token that you send via language model.

And it is often becomes larger, than the weight of the model itself . Do you have these? crystallized knowledge in model weights , and you have dynamic knowledge of the exact the conversation we had we are leading, in the KB cache, as I I like to think about it. Yes. And that's why sometimes people are watching, how deep in conversation everything begins to deteriorate due to some technical problem. Yes. So the KB cache is enough interesting in this in relation. KB cache is accurate representation everything that was earlier.

We keep all the information that seen in the conversation. However, during learning model studied mainly not for very long contextual dialogues . She studied mostly, let's say, on 8000 or 16000 dialogues with tokens. So, if take a model up to 200,000 tokens, then certain training took place and at this length context, but it is not main force models. Therefore, for Frontier Labs have always been the task of finding out, how to make a model just as much intellectual at 10,000 tokens, like us we expect at 200,000 tokens.

And this will be eternal struggle for us. We already have a lot years there is a concept 1 a million contextual windows. Enthropic, it seems , was the first to reached the length context window in 1 million. I still am I use/ I compact in my cloud code long before the length context of 1 million. I don't think so. really great to achieve complete length. And therefore these extremely quick approaches with extremely low delay in the end limited by this factor. Yes, you can do anything, from scales.

Quite possible to have unsurpassed performance at storage of scales. However knowledge base cache will be a big obstacle. So after three years, in five years, which role, in your opinion, will play these types of chips? Which market share they have on heterogeneous market chips? Crisis, Grock and maybe more a few others, you should think about them as about accelerators. They really good at used together with a more traditional a device similar to to graphic processor, which critically important has built-in off-chip memory. You need off-chip memory for capacity and built-in memory for speed.

We want hybridize these two things. So, if you will you take transformers in within the limits, you will translate transformer to a million contexts length. Ultimately, in you have this one, you know, stage associated with calculations, which consists in the actual matrix multiplication for what we call MLP, where encoded most of the knowledge world of knowledge models, and then you have a level attention, where are we dynamically we adapt to current conversation. Attention to the limit usually associated with memory, and MLP is in the limit related to calculations at big enough package size.

And I would said that original sin transformers is that you took this one extremely fundamentally related to memory level and placed its next to the level, related to calculations. Very it's hard to have one chip, which performs well as computational operations, and memory operations. Graphics processor quite balanced in this regard, but you will have to choose one or the other . Cerebrus has a very quick access to memory for something on like multiplication matrices, and very good to place MLP, by essence, weight, on a chip Cerebrus.

But graphic the processor has ability scale to very long lengths context, so you would like to focus attention, uh, maybe on graphics processor and MLP on the Cerebrus chip. And I think that's exactly what happening with Nvidia and Grock. Can you to talk for a minute about transformers? THERE ARE -uh, yes, you are so good explained some basic concepts for people who, again, not very good though familiar with this innovation 2017 year, for example, what are his strong and weak parties, and whether Do you think he is?

will remain dominant architecture for future AI. It allowed us to very to study effectively unsupervised data, because transformers, in the end, they take any what sequence, any arbitrary data sequence, and are trying to find patterns in these data, and they critically affect for the operation of attention, which is the main, uh, component transformers. It allows the model dynamically to adapt to what does she think most important component sequences. WITH with every step that you do through transformer, you, by In essence, rethink input data that viewed previously, and determine which of them most relevant for your next forecast.

So this is extremely light to study, uh, arbitrary data sequences. The most interesting data sequence, which we regularly we create, there is language, and that's how we achieved dominance in language mode. But, uh -eh, if even more to deviate, I think, what transformers really good coped with scaling. Transformers are not create such human a priori . Transformers just they say: "Well, in data sequences will be certain regularity, and if the pattern Yes, I will find her. I I am going use everything more and more parameters for solving this problems while she "It won't work." Uh, and transformers also receive benefit from a large amount of work computer vision.

For example, one of problems in computer vision was what we had it is difficult to move on from hundreds of thousands the parameters you you get for linear models, such as the method support vectors or others are outdated machine model teaching. They had, you know, thousands parameters. Then we moved to deep learning and received dozens millions of parameters by means of computer vision. The largest models were, you know, close 150 million parameters that were a huge model for computer vision. And now we we speak regularly about trillions parameters, and transformers are a connecting link for transition from millions to trillions parameters.

So, if I think about important units scaling-data and calculation, whether there is reason to think that transformers will just stay, because that's what we we are good at it get more from these two things? Well, data is open question, but calculations...Yes, transformers are simply wonderful sponges, you know, you increase computational opportunities available transformer, in 10 times, and you will get some improvement logarithms somewhere. Uh, and so far the laws scaling really are working. They really quite good. And to the point that, I think, transformers Are they doing really well?

They are spreading almost any the data set you You can throw it at them. They are extremely powerful for general education. And I think that transformers especially useful compared to others the methods we use tried to replace attention, so this is what they represent any even number the connection you have you want. Any token in sequence can to be associated with any other token in sequences. So, if in sequence is there any at all connection, you will find him with the help of transformer. Maybe you don't modeling required "everything to everyone".

You don't everyone needs to token considered every other token. But if you need , transformers provide you with such possibility. And while we we don't know any better way to reduce this space, better way to do informational modeling more selective, attention is very, very good operation. This is another one the kind of trick we studied during the time computer vision. Like one of the old ones of the sayings of Carpathia, you Do you know if you have any? new data set, for which you want train the model.

Yours the first goal should be excessive parameterization models and attempt reconfigure available data to to prove that there is the connection you have you can simulate or remember that your algorithm learning works, uh, what can you do to implement knowledge in model. As soon as you you will be able reconfigure, you you can squeeze, and compression is how you you receive generalization. You are not do you actually want remember data, which you have in front of you . Do you want to generalize, and therefore as soon as you reconfigure the set data, you can to work in in the opposite direction and try to find general patterns that fit into the smallest possible number of parameters .

What is your forecast for the future of data and the importance of data in this whole story? I like it phrase that the Internet was once a subsidy on the data. We received them for free. Uh, this is extraordinary high quality, about 30 trillion tokens high quality text. Uh, 300 trillions of tokens, if you look more broadly on the fact that qualifies as good text. And we basically everything This has been considered. Models many times already seen all over the Internet. And with data about people from Internet more nothing is possible make.

Following data processing stage, in my opinion, it is self-improvement models through gyms RL environment. Essentially, we don't even we benefit from random interactions users from artificial intelligence. Before, you know, new the type of data that we very interested, were interaction data people who use chatbot and send signals about tet- signals that they like it or not. Now I like the argument , that the median model, which we serve, much more advanced than the one that is random the person gives the opposite connection that the signal you receive from random human preference, or, I I think, unconditionally human preference, actually more worthless.

On at this moment you expert advice needed human preferences. The model has grown everyday generalization. Yes. Daily, Joe. Exactly so . So the feature data for me is to model dates are difficult verified task and let her to work in this the gym, where she is in a certain way isolated, and she has there is just a problem, over which she can to advance, and get measurements whether she has achieved progress in this problem or not. You can you imagine that problems with coding belongs to to this category.

Mathematical problems also belong to of this category. And, hm, more and more often we can just give it to the agent computer, and it behaves like a human being- employee, and simply give him feedback about whether he achieves progress in achieving the target result. It the environment becomes data. I think that this not too much differentiated approach, but, um, it was very, very productive because What have I seen so far? And you think that this just continues during a very long period time? Is this another one?

example, if I think about the Internet as a one big block like this one, another big one block that will have your, you know, day on the sun, and we all seem to be we will get this, and then we will have to to move on to something another. I think that this actually deeper. On Essentially, the idea is to because if you need a general artificial intelligence, the best way to achieve it is just continue accumulate specialized intellects, as long as you have there will be no more left gaps for filling.

And the test here, the only thing that you need to do, for this to work, this make sure that your the task can be verify. You need to provide model system self-assessment. If you have it, you have it recipe self-improvement in any task, which one do you like? And I think you saw, how is this is confirmed by the fact that how they spend money advanced laboratories . They used to spent a lot more on data. Now they spend much more on RL environments, and these environment completely reflect this recursive connection self-improvement in task that can be verify.

Good. Me I like that we have deviated from the topic on various small underdogs, but we return to your initial task-to-do existing equipment more effective. Yes. Thanks to the best control over what takes place on at the hardware level, for with help software software. Ahem, therefore, therefore, simply keep doing what you are doing have done so far, and what you want to do, and then we will move on to equipment, and then, finally, to energy. Sounds good. Therefore, yes, I mentioned kernels. It's strange that people think that the nuclei finished.

There are wonderful people like Triau, who write wonderful nuclei, and they form the basis all of ours, hmm, modern deep teaching, built on lightning fast taking into account. Modern transformers built on lightning fast taking into account. But if you at all deviate from good luck if a new model will appear , which has a slightly different embedding method positional information such as change of RoPE system, suddenly the core, which in we were, not suitable for this new model, and us, maybe you will have to make a patch for of this nucleus.

I wouldn't said that we we are at a stage, when we have to invent new kernels from scratch. But the mother opportunity quickly to change the existing GPU core, sorry, core, by the way, it's a general term for any the program you run on GPU. So historically cores, as usually placed to the library, where each core has a very limited area application. Usually you have a kernel for matrix multiplication. You have a different kernel. even for something as simple as addition. Do you want add two tensors together.

This is a different kernel. AND then we more and more often we begin combine these cores. So, if I do matrix multiplication, and then I want to add it to another matrix, which I also multiplied, maybe these two will become one nucleus, and I'll just combine operations. Where instead of in order to record data into DRAM, and then read them back, just to fulfill adding, maybe I I can just do it. That's, uh, easy. Why are people still doing this? are engaged in? It seems what is artificial intelligence would be extremely skilled in development more efficient cores.

Maybe that's where we are we're going, and we're just not quite there. But if we're not there yet, then Are we going there? If we are not there yet, why people still do this are engaged in? Why Tree so famous? You know , that's a name I know. I don't want to talk. on behalf of Tree, but that, What did he teach me? this is what, uh, more not necessarily write kernels by hand. I love to say what we write kernels on the board. We let's go to the board, we describe what, in our opinion, opinion, should do machine.

Then we briefly describe this is natural language for the model, and then the model is capable to do...okay, here are my inbox and source data. Here strategy of how we we want to pass this on work on graphics processor. I am going to implement it. So, we do conceptual design. That's right. And I think I not sure why the models are not great are dealing with this. I I don't think it's like ours. remote control management or something similar. I am sure that in 6 months there will be many of us the best models, uh, with kernel development, and I sure that the labs will say you that they already make a significant part of your kernel development, uh, fully automated way.

And therefore software provision as advantage, Yes, if I think about software provision as a as close as possible to the speed of light, effective using basic hardware software will stop over time be an advantage for such a company as your. This is true. Growing a wave of something on like Mythos or GBT 5.6 Soul which raises all boats. That's true. Hmm, I don't really I think it makes sense. specialize in in order to say that we are working on improving the model for engineering only kernels.

I think that this not really the most significant subset encoding in general, uh, core engineering in particular. Maybe there is some information about the privilege you have insert into invitation, what is it in a useful way direct the model to better writing cores, but in general saying yes, we all we are below the limit from the point of view of these opportunities. Me always like this is, uh, from history energy, there is always such a pendulum between raw materials, let's say, coal, and then, if there is a certain amount of energy, accessible inside a piece of coal, for example, which one percentage of it we we can use, and most of energy history consisted in the fact that get this number from 10% to 95% or how much whatever, where we are, for example, if I I'll just think about it.

in Blackwell or something this, and Blackwell is this a piece of coal, yes, on what percentage, on in your opinion, we we are, for example, how much effectively we can use existing piece today? There are many different methods of analysis this. I think that on to a certain extent we really effective optimize productivity when graphics processor does what he has to most like, namely applies the matrix large size, which performs work on peak level 70-80% load, and this is not limited software provision, and power. How Nvidia indicates peak bandwidth, somewhat optimistic.

You you will never reach this because power limitation , ah, hmm, due to heating. Yes, that's right. Thermal regime, let's say 70-80%, he saturated, that's enough good, but in practice you don't spend more part of his time in transformer for on that happy path, where do you perform a large party matrix multiplication m, ago, uh, our job is to build an engine around the chip like this so that we constantly downloaded graphics processor these big ones the amount of work, and one of the deepest transitions that we had in the world graphics processors over the last year, was this transition, you know, you no longer program one graphics processor at a time, you should think about the whole rack , and maybe you should to think about the whole cluster, the whole center data processing at a time.

And again with Nvidia, they started to supply not only one graphic processor or one motherboard, and actually the entire rack the system they are appointed. They They call it NVL 72. Uh, their last chip, Grace Blackwell 300, supplied in in the form of a rack for 72 units. And this open race to find out who can program the whole rack computer scale as possible more efficient. And I I think that this one type of calculations–this future as efficiency, and speed. Actually, Nvidia is great copes with the fact that what if you want as low as possible delay, you should use this chip.

And if you want as high as possible bandwidth, you probably do too should be used this chip is already there. And it all comes down to that this is very new paradigm programming. One of the things you you hear, is that, that the market of the best chips, say, Blackwell, now it looks like a market drugs or something like this. As if a lot is happening exciting things to get them as soon as possible more, because everyone so short. I I would like you to reacted to this analogy, for example, is that feeling but also talked about what the market is like, for example, not the most modern chips, for example, if I ready to accept a little or moderately worse chip, What is this market?

For example, let in us into this world. Yes. Good. Sprat things. First of all, uh, yes, mostly Nvidia has a long-term a look at all yours chips. They see this huge demand for black chipped walls, and they can do what others did suppliers in past, that is just raise prices and satisfy market needs. You know, demand curves and offers will be corrected; they will intersect at a certain moment, and everyone will be technically happier . But Nvidia sees that if they just will allow the richest to buy all chips, it can to harm them in long-term in the future, if this the client ultimately will get much more power.

They understand that computing is power today, because, uh, they are enough strategically suitable for distribution computational resources. This is the first opinion. Second opinion is that relationships have great importance. Nobody wants to receive a huge rental order chips from this a new startup that says: "Oh yes, I I am going to rent. 10,000 black wells for 3 or 5 years". Startup usually only works several months. Who knows if they are good for your money. Way to convince someone give you access to calculations nowadays quite complex and requires or excellent relationships, or, uh, simply incredible financial support to make it happen with Nvidia side, and all that because the deficit so tall, and demand is simple going off scale.

Regarding others chips, I wouldn't even called them worse. I I like to say that there are no bad chips. There are really bad prices. And, uh, I'll do this, so that any chip worked for at the right price. It one of a kind company ethos. AND let's talk about AMD. AMD, I think, great chips overall. The problem is because people are not very understand well how program them. So, you know, I told you about that , how is ours so wonderful development team kernels. They are so are taken seriously before squeezing productivity with equipment.

Nvidia does quite well this is for your own chips. Honestly, there is a certain alpha version, which we can to extract, but actually there is much more work on other chips, so that the supplier performs a little less work than Nvidia to create the best kernels "out of the box" or even better for me . There is an alpha version, and, as and in other people, there is the idea that AMD not as good as Nvidia . This is music for my ears. I am very glad that they sleep on this Chipi, I'll buy it.

as much as I can . Now, I think this is actually not very much anymore truth. I think AMD actually, uh, something popular among some big ones buyers. Um, you know, I think it's public. Meta and OpenAI bought a bunch AMD chips, so we all we see more often that, uh- e, all AMD deliveries are also distributed . But there is a long story other companies that appear; but, new companies are great, uh, such as Etched, or Senova, or dmatrix. All these companies appear, and I think the main thing the challenge for them is scale.

Will they be able to they get from TSMC enough plates, to produce chips and bring them to market? But, of course, if a new one will appear on the market Chip, I would like to to learn about him as soon as possible and assess whether we can buy a significant part of this supply. So for you, in essence, arbitration. For example, if you can much better extract performance with chips received less attention, you you can resell this is with a markup, and this could be great business.

That's right. That's right. And I I think it's not about because everyone else just have skills problems, because of which they are not can force these chips work like this just fine. I think that we are enough competent in this. I think we probably one of the best teams in the world, which uses several silicon architectures and quite aggressively seeks reduce productivity in unexpected places. But, um, I guess what is speed, with which we are ready for build our stack around the new chip. We don't have a big number of officials , our suppliers service centers data processing are limited only to this class of chips, and we it will be very difficult deploy these brand new chips.

We have several very creative partners from processing centers data that is ready to move very quickly forward, and there is a new class those we are talking about we can talk, and most importantly, we do not We are afraid of the challenge. Honestly, big part of this is to just say: yes, we love TPU, we Let's make TPU work. Yes, we love Tranium. We we will force Tranium work, and if he won't work, um, easy, we will find way to fit it in heterogeneous system service.

He will have its place. Each chip has comparative advantage . We have to find this one. advantage, and then use it in in this direction. It just like a break, before we move on to the equipment, processing centers data, energy etc., which will be true interesting part conversations. I would like to so that you can tell me about your perception class anxiety investors. Yes. How do you look at shares of companies that are engaged in memory. Yes. Hmm, or mine favorite option now–this is the schedule, which shows interest semiconductors in S&P 500.

Historically it was about 2 3 4%. Now this 19 20 21%. And it seems that if you study history market, you see everyone these things over time that have reached some crazy short-term peak and then fall to long-term norms. Hmm, and I wonder how this worries everyone investors. Uh, you know, a lot of people earned a lot money on Micron, Skhinx and similar companies, but everyone feels that in the long term perspective computing is a commodity, and, uh, they don't will be quarters or a fifth the entire market world capitalization, that's why they're afraid, and that's the whole situation.

Ahem, everyone admits that there is a huge deficit, but all think that we are let's decide, and these things will return to their usual places on capital markets. Me I wonder what you think. about this narrative. First of all, I'm not so much I study history, How much do you know her? history. I was born in 1997, and my mom worked at Intel in 2000 year, on the eve of 2000 year when it took place boom and bust, and you know, I I remember the times when Cisco was the most valuable company in the world, and Intel was a little behind.

I mean, I I mostly spend parallels with that period of history from 25 years ago to today. And I think, what is the main thing the difference is in that it is significant part of the investment in network equipment historically was speculative. We predicted this future demand for users who never showed up. And I I think it's interesting in consumption of tokens or consumption artificial intelligence in general, what is this no longer speculatively. People buy tokens, so that they are immediately valuable for them. You are not accumulate tokens, you use them immediately.

This is even differs from of what was two years ago because when in 2023-2024 years was observed shortage of chips for generation of bunkers. THERE ARE -uh, at that time everyone the expenses were focused on learning, and learning is essentially speculative. Now everyone installs restrictions on as much as possible spend on cloud code. This is completely different. world—talk about logical costs conclusion and predict them growth. I think that logical costs conclusion monotonously are growing. Uh, no. no speculation regarding expenses for logical conclusion. Vanta automates safety and conformity requirements for over 16,000 fast-growing companies like Ramp, Cursor and Harvey, around the clock supporting their audit readiness.

This is the Agentic Trust platform number one, and now she helps such to companies like yours, monitor risks, which arise during audits of your suppliers, your tools artificial intelligence and all yours environment. Everyone new tool, on who signs your team, everyone supplier who enables functions artificial intelligence, —this is an opportunity for for something to go wrong not so. And most security programs are not were created for the growth rate of AI. Agent Vant works as GRC engineer 24/7 in the background, finding problems, developing for you correction and reducing time assessments suppliers up to 50%.

Regardless of whether are you fast-growing startup or global company, Vantaa will help you to conquer and prove trust. Invest, as the best listeners. Get a special offer from $1000 discount USA from Vanta on vanta.com/invest. Ridgeline is first comprehensive accounting system with built-in artificial intellect for investment companies, which deals with accounting briefcase, agreement, reporting, trade and compliance with the requirements, all on one unified platform. Companies moving from obsolete technologies to Ridgeline because of how much functions of artificial Ridgeline intelligence ahead of anything other in software provision for management investments.

I heard from many investment managers about artificial intelligence, and they are approximately are divided into two camps: some not know what to start at all, but others are convinced that can create own system management total orders for the weekend. The reality is that management investment firm will always require management, control and a single source reliable information for your data. And none enthusiasm for artificial intelligence does not change this requirement . If you are serious relate to your strategy companies regarding artificial intelligence, Ridgeline has to be part of this conversations.

You can order demonstration on ridgeline.ai. Returning to your view on hardware software. Me interested in the level units, for example, talking about chips, racks, clusters. I would wanted to talk about data centers, and you said that you have there were interesting partners, who did cool things things. Tell us about the present and the future of centers data processing, as you You see it. Because it seems you you know what's critical important thing for in order to be able to to serve the whole this conclusion is many innovations in this part of the world, and, obviously you focused on this.

I think that one of the things our conversation returned to that, what is training against the conclusion, for example, which one the difference between these two categories, and you know what two differed years ago orientation for training from orientation to conclusion today? AND I think that the most conservative players in everything AI stacks should be players in infrastructure, be then processing centers data or even more conservative, TSMC, people involved infrastructure with chips. Uh, and that's why data centers historically built as processing centers artificial intelligence data intelligence; they were built for learning, uh.

AND learning is superplural working load on logical conclusion; you can force any educational cluster work for logical conclusion, but, maybe not the other way around. AND What is the difference in networks? : uh, how old are you invest in bandwidth between chips, and which one cluster size you Need it? Actually there is a certain inefficiency scaling for processing centers data. For example, much more expensive and more difficult to build , you know, 100,000 graphics processors in one center data processing than build 10,000 than build 1000.

And we now just we are talking about, How many megawatts do you have? or gigawatt? And in basically no easy way build gigawatt center data processing in United States, e- e. Even 100 megawatts is becoming more and more more difficult. It practically impossible, if you are not very special group customers. Uh, 10 megawatt–this is probably on the edge of the possible today, and 1 megawatt, I would say, already enough. So on market exists incredible information about what you can find a lot total capacity, but she won't concentrated, and this no one was interested who builds educational data centers, because you just assume that everything will be in one place, because no one wants to have deal with intercenter training.

So the market is somewhat lags behind. I think that the market is still assumes that we will have to get rid of these processing centers data per 100 megawatts and 10 megawatts, wherever we put them didn't find it, that's all another position that I I hear from many center developers data processing, but more and more often we we see how some new thinkers realize that logical conclusion will be suitable for these distributed processing centers data with power of 1 megawatt, and, er, we completely agree with this, and we are very happy to buy small pools calculations throughout United territories States and use them like our park of logic calculations.

This gives us idea of literal physical the size of, uh, 1 megawatt against 10. Yes. Well, it became really strange with the appearance of liquid cooling. Now you can you fit incredible level power density into one physical rack . For example, megawatt calculations, you can imagine it as huge hall data processing, as huge warehouse. AND now you can actually fit it in...Yes, approximately eight racks calculations. Each rack size approximately from refrigerator. You you can just imagine eight of them, lined up. Hmm, yes, that's a megawatt.

AND so your opinion is that the future that you want to help to build, is a whole a bunch of different chips, which can be use together. Yes. Which ones can you buy? you know, you are a buyer, to get maximum per chip. And what about these chips then? can be combined into very small processing centers data, to just do conclusion. And what are these two a whole bunch of steps random calculations , some of which are cheaper than they should be, your ability consume more of them, and then small units of expression in data center are much equal cheaper intelligence.

I definitely think so. Yes , there are many ways access to cheaper failures, if you are able show creativity in their decisions. And therefore, one of the ways in which I I describe what we we do is to because we buy any chip in any where in the world is any period of time . This is the level of flexibility and liquidity, which , I think, not now who is not there We are very hard we are concerned about to invest money to where we are talking, and we will really take any capacity and we will find a way, so that they work in our park equipment.

And this most of our advantages today and in the long term perspective. To us had to create more than this advantage, investing in these data centers, to which other people will be skeptical to treat because you you know what will happen, when you create this an army of a thousand small centers data processing vs. one, one large gigawatt data center. Well, a few things. IN you will not be there backup power quite often. You don't have there will be backups diesel generators on site. All this is very expensive.

We we reduce all these overhead. IN in many cases in we won't even be there network redundancy. We will place them on facilities where we have good access to power supply, single power source, and we will lay one line fiber optic cable to these centers data processing, but in there won't be three of us lines fiber optic cable with reservation, reservation and agreements on the level service. Sometimes it's simple will fail. Actually, I don't I would be surprised if some of them will receive 95% trouble-free operation, which is bad.

Very bad. It fatal, terrible for anyone else. If you survive in big...in you will be practically absent buyers per center data processing with 95% trouble-free operation. I am the first buyer. , I will buy 95% trouble-free operation, and the reason for this consists in this background mechanism, what if something works in background, you indifferently. Partly, uh, this actually two things. First, we have really reliable control plane, which is wonderful will cope with any what separate failure in any individual data center, if he doesn't connected with others processing centers data, and I can just to postpone work loading somewhere yet; I'm fine with that.

Hmm, glitches happen. with a certain frequency, and I mostly linear satisfied with the center data processing, which has 95% uptime work vs 98% vs 99%. It's just linear. good or bad for me. Hmm. Now you need that asynchronous part, which I mentioned, you know, we we serve these agents with long horizon, because when the request is not I can do it. will have to look for new graphic processor to place this request. And this means that during this one-step work agent he works hour, but then encounters obstacle, because its graphic processor turns off.

In this this agent at this point in time will perhaps experience extra minute, two or three, maybe even 10 minutes of delay. But my argument is in that my customers don't care because their agent worked for hours. They are sleeping. It doesn't matter. Not it matters if one turn sometimes becomes a little longer. Yes. So we say to our customers who our average capacity will be very competitive , but our P99, our delay of the 99th percentile, there will be no be controlled, its impossible, and I will give you instead.

unsurpassed economy, and I think that this fits right for background, agents are talking about power as the category that you You see, this is interesting, innovatively, where, on In your opinion, this passes okay, so I said I want 95% of the time trouble-free operation in my centers data processing, or I can even get 80% of the time trouble-free operation at the right price, probably, um, and what is it means, well, I'm the son California, I love solar and wind energy. I think that is solar and windy energy not enough used in United States, and always a problem was this instability.

You would even considered solar and wind energy unsuitable for processing centers data because you have constant base load and intermittent source feeding. What are you Are you going to do it? Well , I think we not really far from a solution this problem. I quite capable withstand disconnection electricity in my processing center data, which measured in pairs days or weeks that is the worst a nightmare scenario for the machining center data; will we have long-term disconnection electricity, therefore that the wind is blowing, and the clouds in the sky?

The fog hangs above the valley some time. This is the worst. scenario. Actually he is very predictable, and I I can just to call forth powers in some other place in the world when it will happen. Uh, me I'll just model it. the weather and I'll find out when my processing center data will be disconnected, I will transfer my data and workload somewhere else, and everything will be fine good. The secret is, What will this give me? better access to power supply to which no one else will touch, because It's so annoying to have a mother.

deal with such disconnection. And if my chips quite cheap, they probably don't will be Nvidia racks. And if my chips cheap enough, I not against capital costs of availability non-working chips. Yes. So, I heard you describe all this system as a strategy garbage collection. Yes, That's right. Yes. Uncover this analogy a little. Well, first we collect chips, and then we are gathering capacity for these chips. Idea is that in in both cases I don't I want to compete with anthropogenic or open artificial intelligence by computational power.

I am not I'm going to win. them, and I don't want to. I I want to be more creative and use the offer they today they don't count understandable. And over time I will accumulate enough aggregate offer. I I will never get it. concentrated proposal. I will get only cumulative proposal. And over time I will build mine. aggregate plant, which is unbeatable from an economic point of view vision. We are building a factory. We we are trying build the best steel mill in the world. Uh, but he will be produced through mini-factories, and not because of big monolithic steel mills.

And if I imagine different versions of this, for example, how much vertically integrated can be be... Yes. One version would be extreme when you You own everything. So this is very capital-intensive business. You own a source known to you energy. You are building data centers. You are developing own chips. You you control software provision that consumes the most of these chips. And you selling the final token to your user as if your user —it's me, and you just own the whole stack. But you can imagine many others business options, where do you know that can you draw a line anywhere.

You can to be incredible easy capital, to own nothing and just to be coordinating plane for all this. Virtual collector garbage, or wrong? What do you think about this? question, what type business to be? You know, there actually is two parts of me that answer this question. One— CEO company, which need to work every day and develop as stable as possible and as soon as possible. Other —founder, and the founder is much more inventive and just loves these things . The founder in me wants to do everything.

It my whole life. I am everything I spent my life thinking about chips, energy, power. Me only this worries me. So, Of course, I want to be most ambitious . I don't want to stop. never. I never I will stop until I will build the most effective system from soup to nuts. You are, in essence, use factorial of real life. Very much. Very strongly. So this is emotional flow a complicated answer. WITH general's side director i think we need to be more pragmatic. I think that capital, which we are considering for possession of all, as you said, this madness.

Yes, software provision has high leverage, therefore we must start with software provision, but after all, you know, whether we own production electricity, or can we conclude great deals on buying and selling electricity from communal enterprises? I more inclined to in order to allow to other people specialize in because of what they are historically good understand, and then to see if can we achieve scale. I am thinking about it's like I want to to achieve scale, when will I get the right take it upon yourself wing.

I absolutely sure that efficiency can to be achieved everywhere. If you you can destroy assumption that people, in which ones will I buy, they did assumption that who will be theirs customers. And I, maybe I'll destroy these assumption. This is enough. optimistic view . Uh, I think that's it. possible only because we really are we are trying to support largest market calculations in history computational techniques. We actually we are going invest billions , trillion dollars in logical conclusion. THERE ARE- uh, and thanks to this focus makes sense create a lot things that are individual for logical conclusion.

AND my job is to in order to search for all places where it is possible. And then, as soon as they will become obvious to me, my partners and my partners, I I will ask my friends. partners to create for me individual things. AND if they can't do it for me, I I'll do it myself. If you had to just move away from all this systems, software software, equipment, energy etc., and rank places where, on your opinion, we today the most ineffective in production of useful intellectual tokens...

Yes. What does this look like? list? I think that scaling calculations actually very effective. Uh, for example, if you give me more flops, I will use more flops. And I would said that we already quite sensibly we use flops . Uh, if you look on a modern model, then very few models have density more than 10%, i.e. activated 10%off possible amount experts who can be activate. And I think, that Frontier models are closer up to 1%. So they already quite sparse. I I don't think we we spend too much resources on the side MOE.

People work with they have been around for quite some time . They are quite good. cope with compression. Where we are not very good we're coping, that's it with attention and her using memory. In particular, cache KV is enough now uncompressed. I think what if you look at entropy in the KV cache, then she is not at all earns a living. We we save a lot kilobytes of cached data KV per token. Hmm, and this, probably by an order of magnitude or two. And I don't know, What does Frontier do?

Labs, but Deep Seek, definitely publishes very interesting work, to compress it further and so on. And they achieve good progress. And I think, that the fact that they able to achieve progress by an order of magnitude every year or so moreover, signals about that there is still much opportunities for development. I think that all of this is on a micro level. If even more to move away, I think, that we really don't effectively we distribute our calculation. For example, this year we spent all these calculations by 5 millions of Blackwell chips .

Where are they all going? Are they all? are used constantly? I, of course, I doubt it. I think that at a certain level we just need better orchestration calculations throughout world. Uh, that's very difficult to do, so that's a lot of calculations disappears into private computing pools that will never see light of day, and these graphics processors very sad idle. Uh, actually me physically painful to see that these graphics processors- it's simple, you know, silicon and energy, which went to them, and they just idle, and I want to fix this.

How we organize and orchestrate global computing as a shared resource and we pack them more efficient. I would assumed that, you know, we all laugh at XAI for what it has certain problems with general using the flop on their clusters, but the reality for the rest of the world is such that everything is much worse. A bunch of graphics processors simply is in storage or in private pools, allocated to a specific client, and they just don't are used. You directly attack the effectiveness of this. This is much more efficient.

Yes. And how? about factories? How are you? What do you think the future holds? the factories themselves? I think everyone is asking the question of whether they can manufacturing companies memory, TSMC, or Intel and others will be able to expand your power? Yes. Or we'll do it here, in USA? Yes. Production the Riffon chips themselves. If we could just snap your fingers and to have 100 times more chips in stock today, here, probably would have been much cheaper tokens. So yes, that's it. seems important part of the universe, on which is worth hearing from you opinion.

Yes. Well, that's interesting. All growing in the balance with each other, right? Yes? If we click fingers and double all these things, you can to solve a narrow place TSMC and you just you will encounter another one a bottleneck. If produce 20% more chips, then you immediately have another one appears bottleneck. However, I I will say that it is interesting that what do they think mandatory supply. Uh, that, what do they think invariant, which their customers always will want, on unlike what I I think more flexible relationships .

I think that if factory opens I have more of my own. compromises, I can take more smart decisions about what, in my opinion, I can do it. One of the most interesting examples here is that any factory has big gap between the worst chip that comes out of production line, and the best chip that comes out of production line. There is a big difference. in how chips are made. Uh, and then the question is is that if you have one company like TSMC, they very persistently are working to to reinforce what we we call these process angles.

We we want the worst The chip was as closer than characteristics to the best, and they receive a wonderful lens to make it happen possible. But this means that they add a lot controls in the process, which may , I don't need them. Maybe I really am ready to find a place for this worst chip. You don't need to so much to strengthen control process that takes more time and money . Uh, maybe I am. ready to accept much more defective products . And I think that for us it's like a more holistic optimization around the cost of crystals, supply of crystals , and then the cost power supply and places where we can them to place.

And all mine the goal is to so that so sharply expand supply power supply throughout United territories States, so that I have home for many chips that are otherwise not would deserve theirs places in the center data processing. Can we to talk about how you developed a system your own business? What lessons have you learned? ? You talked about some interesting lessons Nvidia. Yes. But tell me. about culture and about the way you are structure the team and business where it is the north star. I think that a lot people, you know, in Limits think that we we don't worry about direct character, for example, when do we start to work on model, efficiency will not be very good But we thinking about where we are we may find ourselves approximately in a month, six months or a year.

We are not perceive the state cars on which we work as corrected, even something like a chip Blackwell, if we consider, that there is some kind of narrow a place that interferes we can achieve this productivity. I have meaning that for me It is very important that we well understood and characterized it and recorded so that we could tell about is it Nvidia or friends, and also remember about this is for the future chips that we buy. We want to study things that, in our opinion, opinion, is, in fact, unchanging for us or companies, uh, in long-term perspective, and take this into account future decisions, which we accept.

We very prone to cooperation. I think that one of the most important the traits we are looking for - these are people who are both good students and wonderful teachers. Hmm, many people in our the team had assistants teachers in college, and they really liked it exchange experience knowledge in this way. We are constantly we hold sessions near boards, and I think that collegial an atmosphere where everyone has something to teach and what to learn, extremely important for us. What qualities the people you wanted would hire, on your opinion, will be resistant to work Wednesdays in 3 years when more things will be processed machinery?

Curiosity. This is 100% curiosity. You know, the only thing I can't do to teach, it is love for productivity, love for the sake of to delve into every microscopy, with how the machine works, and understand that takes place on car at this time. For This is the most important thing to me. trait for an engineer with productivity, and this what I'm looking for. I am not I am looking for a lot of experience. working with artificial intellect. I generally I'm not looking for experience. working with CUDA.

It really huge diversionary maneuver . I mean, CUDA as a concept or GP as concept strongly evolved over the last 5 years. None sense to demand 10 years of experience. I want teach this, but I I can't teach. love for productive engineering. This is it. I am eager. Can you estimate? main laboratories one after another, and also closed loop connection code as a category with open source and that what do you think, happens and will take place successively? I would said that laboratories pay a huge prize to be for 3-6 months ahead of everything another.

Uh, and I think, that this is probably still It's worth it. I think that for open and anthropic laboratories entirely it is logical to do what they do. You know, there is a sensitive topic around distillation, which, in my opinion, is key element the relationship between closed and open frontier. And you know, I would like to offer alternative view of this, namely, that there is a feeling that distillation is theft that you take something from the models frontier, when distill based on their results. AND in fact, even if that is not your intention, even if you never don't try, you know, collect data from anthropogenic sources, one thing I would suggest, this is what is increasingly higher percentage artifacts that we we place in Internet, created artificial intelligence.

Even if you just just look at GitHub, you know what? interest repositories, created by last year, on our opinion, was created cloud code? Hmm, is it we think so distillation? Because That's probably all we need. necessary. I am not I would be surprised if you you can teach Fable Glass model only on code results, which one do you think good on GitHub with open weekend code. And, of course, if we accept position that users possess the results of its interactions with artificial intelligence, and they decide to place it's on GitHub, which is a lot of They do it, we we will have latent distillation during for a long time.

Me it seems fundamentally impossible. Like me, I I don't think so. in principle possible prevent the spread information or model capabilities. It will happen. Question only in how fast . And then there is question: are the laws about scaling and improvements are in effect forever or throughout very long period of time? And if yes, there is value in in order to be on three- six months ahead, and this will continue as much as will continue, and they can charge a huge prize for these tokens compared to very cheap token open source.

Or this is the right way thinking? I think it's possible. . I don't know if the award 3-6 months in advance will last so long. I mean, if look at deployment on enterprises, uh, they don't move with at a rate of 3-6 months. Many businesses, probably still use something like 46, Opus 46 or Opus 47. They are not are implementing advanced technologies quickly. In people there is a lot questions about implementation of any What changes at all? And I I think we're just We start so early, which, um, I don't think there is some way to name the winner in these races, and, Of course, I don't even I think these races there may be times somehow solved.

Always exists continuous process, and I don't think so open source once will disappear. If there is vacuum because of the fact that one leader goes, on his place is coming new leader. Too many incentives and too much passing wind. Just from every day becomes easier to train the model frontier-class. So your hope is the future lies in so what is the balance between closed and open, Do you know what balance is? between model companies that they do everything because have the advantage stack ownership or something like that.

You know, Enthropic can do it, you know, it's like new Google, which Google can just do it or something like that. How are you? hope that what will the future look like ? I want a lot of tokens and various resources. I want to everyone created their own own resource, and every company, every company, every user even made the agent his own own. Uh, I think that we are not that far away from this level settings and opportunities. I want, so that people can own with my intellect, and I I want this one intelligence was tuned, probably not because fine tuning weights, and probably through context teaching.

This is more technical detail. But the main contribution to this There is a bright future, after all. essentially, cheap tokens. My job is to in order to do tokens as much as possible cheaper. I will achieve this, and I will do it through each level in stack available to me . I like them. levers of supply. I will use each chip. I will use every source energy and each a piece of land in United States, which, as you know, suitable for this. And instead people will have an incentive to investigate what is to have abundant intelligence.

We still treat agent as a person, consultation with which are expensive, and you must ask him when you have a difficult question. It wrong way of thinking about intelligence. It's unbelievable that the car can think, and we must try to convey this as much as possible more quantity people. You are sitting in such a a unique place and do you have one a unique perspective on what you are trying to do to do, to do this function reality. Which, on your opinion, yours the most divergent worldviews compared to yours friends who are really well informed and interested in this?

For example, what forces your ideas to force your friends to watch you as if you had three heads? Most ideas about chips, I would say. You know, when I say about the creation own chips, and they they ask me: "Oh, so "What has changed?". By In essence, it is about to get around the HPM shortage and focus on more extreme unloading at other forms of memory, such as flash memory. Hmm, I'm very excited. this idea. Everyone in mine the team knows that I constantly hitting drum, for example, what do we need change to architecture of the model, to unload KB cache on flash memory worked on much better equal.

And I do it all the time. I am discussing. It's like in community of people who engaged in logical conclusions. You know, in we have different views on what is possible to do if develop a system, which serves, you know, from one to ten tokens per a second that is ours main star. IN in a broader sense, I think there is one. a broader sense, you know, What are you doing? How people consume a trillion tokens per day? Well, this is the world we live in we want to create opportunities for this.

What is a trillion? tokens, as we know, how much is it a trillion tokens? Well, well, at the price of OpenI, it is at least 5 millions of dollars, at least 5.5 or $5.6 million. Yes, I think so. dollars were probably the best indicator . Yes. Yes. That's millions. dollars. Yes. So, in which in the world we consume what is now worth 5 million dollars for a person per day? Yes. I mean that we asked at least, hmm, you know, cost improvement three-six token orders. Uh, include this in 5000.

In There are probably some of you. clients. And in fact, I would say that we working on a model of a certain size, we approaching trillion tokens, which measured, you know, tens of thousands dollars. And this is what you can you imagine for one task. You at all Are you worried that average person simply cannot and will not wants to do this, for example, does not this is now mine with your own brain? It seems that in reality There is no such thing in the world.

high demand for intelligence. I never did that. I believe. In the world there is always a demand for intelligence. I think that the path and ways to it intelligence is ours challenge as communities developers products. I am not product specialist, so I can't say, who had the best vision. Do you want to give this people an opportunity. I want to give these people possibility. I want, so that they never restrained feeling, what, oh, me, my users free level cannot to use, or I cannot allow it.

to give them so much tokens, and I constantly I hear this from my friends. customers. We want fix it. AND how about the reverse question? Not what you do you think the craziest, and what is the consensus on your opinion, wrong. One of the things to I constantly I'm coming back, this is question about Nvidia. I optimistically set on Nvidia in short-term perspective. AND, you know, about Nvidia, never worth it to bet against them. They always will invent yourself again. But, fundamentally, I think, what is one thing that surprises people, that's when I say them, that if you compare performance 5- nanometer chip Hopper-Blackwell-Reuben per watt, multiplied by Bflat 16, then she is not that strong improved; or even if you do one more step, contact TSMC: if you compare 5- TSMC nanometer chips with four, three and two, then performance on What are these chips?

changes radically. So the consequence of this is what people lose mind through geopolitics, for example, that will happen if we for some reason we will lose access to DMC. And, um, my opposite the idea is that that it won't be like that and bad. Supply, certainly, will experience shock. But the best the processes we have in the West, uh, Intel is not so far behind, in worst case, maybe twice as bad performance on watt, and gap

Go deeper