The Two Harvard Dropouts Who raised $800M to take on NVIDIA

When building something radically new, eliminate general-purpose buffer from your design. The entire semiconductor industry optimizes for flexibility across all use cases, but when you know your specific constraint—like AI chips never running below 80°C—you can make targeted changes throughout the s

June 30, 2026 1h 32m
Invest Like The Best

Key Takeaway

When building something radically new, eliminate general-purpose buffer from your design. The entire semiconductor industry optimizes for flexibility across all use cases, but when you know your specific constraint—like AI chips never running below 80°C—you can make targeted changes throughout the system. These 20%, 50%, and 2x improvements compound into systems that are radically better. Question every default assumption in your field.

Episode Overview

Gavin Uberti and Rob Munro, co-founders of Etched, discuss building specialized AI inference chips that outperform GPUs by rethinking fundamental assumptions in semiconductor design. They explain their low-voltage inference technology, cluster-scale memory architecture, and vertical integration strategy—from chips to racks to production—while sharing lessons on recruiting elite talent, moving with extreme velocity, and building for the future of AI where inference becomes the world's largest market.

Key Insights

The Buffer Problem: Rethinking Industry Assumptions

The entire semiconductor and data center industry is built on buffer—every component designed for general-purpose use across IoT, edge, and data center applications. By focusing on a specific use case (AI inference), Etched questioned default configurations like chips needing to run at full speed at 0°C. Since AI data centers never operate in freezing temperatures, eliminating this constraint alone enables significant performance gains that compound throughout the system.

Low-Voltage Inference: The Thermal Bottleneck Solution

GPUs can't add more compute without thermal throttling—more transistors switching means more power and heat, forcing the chip to lower its clock speed. Etched's breakthrough was recognizing that voltage is quadratically proportional to power (2x voltage = 4x power). By running at under half the voltage of other AI chips through novel power delivery mechanisms, they can pack far more compute into the same silicon area without overheating, fundamentally solving the performance ceiling problem.

Cluster-Scale Memory: Treating Many Chips as One

People ask the wrong question about memory bandwidth per chip instead of bandwidth across the entire cluster. Etched built custom interconnects with 5x lower latency than GPUs (significantly under 4,000 nanoseconds point-to-point), allowing chips to use each other's memory as a single pool. This means scaling from 1 to 8 chips actually delivers near-8x improvement in tokens per second, unlike GPUs where high interconnect latency severely limits multi-chip performance gains.

Velocity Through Vertical Integration

Etched is the only startup building both its own chips and its own racks simultaneously. They developed thermal test chips with exact expected hot spots before silicon returned, validated cooling systems by overpressurizing cold plates, and ran 24/7 development cycles across time zones. This extreme parallelization—building the full stack in parallel rather than sequentially—combined with in-house production capabilities, enables dramatically faster iteration and time-to-market.

The Inference Economy: From Handcrafted to Mass Production

Today's token generation resembles Renaissance-era screw making—handcrafted by relatively small, general-purpose systems. Most goods like iPhones achieved economies of scale where wealth doesn't buy a better product; everyone gets the same phone. AI hasn't reached this yet. Whoever can mass-produce tokens at scale with the efficiency of modern manufacturing will democratize access to the best AI models, making quality AI as universally accessible as consumer electronics.

Notable Quotes

"You kind of have to be sick in the head to join our company. You're going to convince your family to move to San Jose for the semiconductor company run by two what 24 year olds now going against the biggest companies in the world with a design that they're saying is not going to be like 10% better, but it's going to be 10 10x better."

— Gavin Uberti

"We know inference is going to be the biggest market in the world. Whoever produces the most tokens is going to be the most valuable company in the world."

— Rob Munro

"I've never seen an AI data center with ice in it. So, you know, we can feel pretty confident that our chips don't need to run at full speed at 0 degrees CC. In fact, like they're never really going to be running below 80° C anyway."

— Gavin Uberti

"Yesterday this feature wasn't there. Today it's here. I go to show my parents uh and I got this like notification being like, you know, you're all out of image credits today. Like you need you need to get a pro plan. And I was like, holy crap. like this is going to change everything and you know like we clearly don't have the infrastructure to serve it."

— Rob Munro

"The best part is no part. I think for us, it's also the best vendor is no vendor. As much as possible, we want to vertically integrate the entire product. Both because we get more performance, but we can move way faster."

— Gavin Uberti

Action Items

  • 1
    Question Default Assumptions in Your Field

    Identify the 'buffer' in your industry—the general-purpose constraints built into tools, processes, or standards that may not apply to your specific use case. Ask 'why isn't this possible?' and dig deeper when the answer references constraints that may no longer be true. Document assumptions that everyone accepts without question and systematically challenge each one.

  • 2
    Optimize for Velocity Through Parallelization

    Instead of sequential development, identify all components of your product and build them simultaneously. Create test environments and validation systems before your main product is ready. Run 24/7 development cycles across time zones if building hardware. The key is eliminating waiting periods where one team is blocked by another.

  • 3
    Vertically Integrate to Move Faster

    Evaluate which components you currently outsource that create dependencies or slow you down. Consider bringing critical path items in-house, even if it seems unconventional. The goal isn't to do everything yourself, but to eliminate bottlenecks where waiting for vendors or coordinating across organizations slows your iteration speed. Build internal capability for the parts that matter most to your competitive advantage.

  • 4
    Recruit Through Proof, Not Pitch

    When recruiting elite talent who are skeptical of your ambitious claims, offer to prove it. Build simulations, create functional prototypes, or demonstrate technical feasibility in concrete ways. Like Etched's early advisor Mark Ross who asked for a white paper and simulation before believing, give experts the evidence they need to become truth-seekers rather than dismissing you on heuristics alone.

  1. Podcasts
  2. Browse
  3. The Two Harvard Dropouts Who raised $800M to take on NVIDIA