CEO Insider logoCEOInsider
InterviewsAnswers
Get FeaturedContact
CEO Insider logoCEOInsider

Exclusive interviews with founders and CEOs sharing insights for business growth.

Join 10,000+ subscribers for weekly insights

Quick Links

  • All Interviews
  • All Answers
  • Favorites
  • Get Featured
  • Contact

Industries

  • SaaS
  • E-commerce
  • FinTech
  • Health Tech
  • All Industries

Revenue

  • Under $10K
  • $10K-$50K
  • $100K-$500K
  • $1M+
  • All Revenue

Business

  • B2B
  • B2C
  • Subscription
  • Marketplace
  • All Business Models

Team Size

  • Solo Founder
  • 1-5 Employees
  • 6-20 Employees
  • 50+ Employees
  • All Team Sizes

© 2026 CEO Insider. All rights reserved.

TermsPrivacySitemap
Back to all interviews
Andrew Feldman, Co-founder & CEO at Cerebras Systems
4.8/5 Rating
Technology
Approx. $72 Million/mo
Approx. $860 Million ARR

Andrew FeldmanCo-founder & CEO

In this interview, Cerebras Systems co-founder and CEO Andrew Feldman details the engineering behind wafer-scale computing and the company's record-setting semiconductor IPO. Feldman explains why traditional GPUs struggle with fast AI inference, discusses memory and packaging bottlenecks, and explores the scaling requirements of agentic workflows. He shares lessons from the startup's early development struggles, analyzes the dilution of Nvidia's CUDA software moat, and outlines Cerebras's massive datacenter partnership with OpenAI.

Andrew Feldman

Andrew Feldman

Co-founder & CEO

Cerebras Systems

Cerebras Systems

cerebras.aiXLinkedIn

Founder Stats

  • Technology
  • Started 2016
  • Approx. $72 Million/mo
  • 850+ team
  • Sunnyvale, California, United States

About Andrew Feldman

Andrew Feldman is the co-founder and CEO of Cerebras Systems, the developer of the wafer-scale engine, the largest chip in computer history. Previously, Feldman co-founded and served as CEO of SeaMicro, a pioneer in energy-efficient microservers acquired by AMD in 2012. An experienced technology executive, he holds a bachelor's and master's degree from Stanford University and a Master of Business Administration from its Graduate School of Business, driving breakthroughs in high-performance hardware and AI cloud services.

Interview

July 29, 2026

1. Why has token speed become the dominant conversation in the AI sector?2. What does speed mean in terms of user experience metrics?3. How does the Netflix evolution analogy apply to the impact of fast inference?4. How do you define the different choices made in the specialized chip landscape?5. What did Nvidia's acquisition of Groq reveal about the GPU architecture's limits?6. What is Broadcom's Jalapeno chip, and how does it fit into OpenAI's strategy?7. What are the three major silicon manufacturing bottlenecks limiting GPU supply?8. Why does Cerebras's wafer-scale engine process AI workloads faster than standard GPUs?9. How does agentic AI drive an massive shortage of traditional CPUs?10. Why is building model architecture directly into silicon design a structural mistake?11. How did the team navigate the early years when the market was not ready?12. Why do standard GPUs struggle with data movement during the decode phase of inference?13. Why did Cerebras choose to build a dinner-plate-sized wafer-scale chip?14. What was the hardest packaging problem you had to solve during early chip building?15. How does Cerebras manage chip defects and yield reliability at the wafer scale?16. Why do you argue that Nvidia's CUDA software platform is no longer a durable moat?17. How does the 750-megawatt data center deal with OpenAI help scale their cloud footprint?18. Why did TSMC agree to modify their manufacturing process for a 30-person startup in 2017?
Q

Why has token speed become the dominant conversation in the AI sector?

Question 1 of 18
Andrew Feldman

AI transitioned from a parlor trick to an active production tool in mid-2025. The minute you use AI in production, speed matters because fast tokens are directly more productive. Users get more done in less time, making fast inference the primary driver of value.

0
Q

What does speed mean in terms of user experience metrics?

Question 2 of 18
Andrew Feldman

The key metric is tokens per second per user, measuring the speed from the first token to the last token in a response. This is especially critical for agentic workloads with multicycle turns where waiting is amplified. Real-time speed allows users to stay longer and tackle harder problems.

0
Q

How does the Netflix evolution analogy apply to the impact of fast inference?

Question 3 of 18
Andrew Feldman

When the internet was slow, Netflix mailed DVDs. When the internet became fast, they did not optimize mailing envelopes; they became a movie studio. Speed completely changes how technology is used, opening up new domains for AI that are impossible under slow dial-up speeds.

0
Q

How do you define the different choices made in the specialized chip landscape?

Question 4 of 18
Andrew Feldman

We transitioned from one CPU doing everything to a multi-silicon landscape. Traditional GPUs like Nvidia and AMD handle generic workloads. Hyperscalers build internal ASICs like Google's TPU and AWS's Trainium, while independent players like Cerebras design chips optimized solely for AI.

0
Q

What did Nvidia's acquisition of Groq reveal about the GPU architecture's limits?

Question 5 of 18
Andrew Feldman

It was a major validation of our vision. The 20 billion USD technology licensing agreement and acqui-hire made clear to the industry that traditional GPU architectures cannot deliver fast inference. It proved that the AI market requires dedicated inference processors.

0
Q

What is Broadcom's Jalapeno chip, and how does it fit into OpenAI's strategy?

Question 6 of 18
Andrew Feldman

Jalapeno is a custom ASIC co-developed by OpenAI and Broadcom to address their massive compute demands. OpenAI understands exponential adoption curves and has been visionary in securing capacity, striking massive compute and memory deals with us and other partners.

0
Q

What are the three major silicon manufacturing bottlenecks limiting GPU supply?

Question 7 of 18
Andrew Feldman

First, HBM memory is sold out, which we avoid by using SRAM. Second, TSMC's CoWoS packaging process is fully booked, which we do not use. Third, TSMC's 3-nanometer foundry space is highly congested, while we build our chips using their less constrained 5-nanometer process.

0
Q

Why does Cerebras's wafer-scale engine process AI workloads faster than standard GPUs?

Question 8 of 18
Andrew Feldman

AI inference is dominated by moving weights from memory to compute. By building a dinner-plate-sized chip stuffed with fast SRAM instead of HBM, we keep the weights directly on the silicon. This allows us to move data to compute two thousand five hundred times faster than a GPU.

0
Q

How does agentic AI drive an massive shortage of traditional CPUs?

Question 9 of 18
Andrew Feldman

Agentic AI does not just provide answers; it takes actions in the digital world like fetching data or placing orders. The AI chip acts as the brain, directing the CPU as the body. This active loop is driving CPU consumption and demand through the roof.

0
Q

Why is building model architecture directly into silicon design a structural mistake?

Question 10 of 18
Andrew Feldman

Models evolve rapidly. If you build a specific model's architecture directly into your hardware circuitry, your chip will have a very short commercial life. We focus on accelerating the underlying mathematical calculations, like sparse linear algebra, which supports any model.

0
Q

How did the team navigate the early years when the market was not ready?

Question 11 of 18
Andrew Feldman

We were honest with our venture capitalists that we were attacking a massive engineering problem to build something fifty times better than a GPU. During our hardest phase, we spent eight million USD a month for eighteen months before successfully manufacturing our first wafer.

0
Q

Why do standard GPUs struggle with data movement during the decode phase of inference?

Question 12 of 18
Andrew Feldman

In graphics, you move data to the GPU, calculate for a long time, and send the result. Inference is the opposite: you move a massive volume of weights from memory to compute to calculate a single word, and then repeat. Traditional HBM memory bandwidth is too slow for this.

0
Q

Why did Cerebras choose to build a dinner-plate-sized wafer-scale chip?

Question 13 of 18
Andrew Feldman

SRAM memory is exceptionally fast but cannot store much data per unit area. To hold the weights of modern AI models on fast SRAM, we had to build a chip that was forty-six thousand square millimeters, which is fifty-eight times larger than a standard GPU.

0
Q

What was the hardest packaging problem you had to solve during early chip building?

Question 14 of 18
Andrew Feldman

It was attaching a wafer-sized chip to a motherboard, supplying it with massive electrical power, and cooling it. No vendors could help us because nothing like it existed. We spent eighteen months failing, analyzing each failure, and inventing new materials to solve it.

0
Q

How does Cerebras manage chip defects and yield reliability at the wafer scale?

Question 15 of 18
Andrew Feldman

Our architecture features a million identical tiles with built-in redundancies. If a single tile fails, we shut it down and route the calculations to a redundant tile. We also water-cool the systems, running them much colder than GPUs to reduce temperature-related failures.

0
Q

Why do you argue that Nvidia's CUDA software platform is no longer a durable moat?

Question 16 of 18
Andrew Feldman

CUDA is losing its dominance. Within a two-year period, seventy percent of state-of-the-art training models shifted away from CUDA, with Gemini training on TPUs and Claude on Trainiums. Furthermore, there is zero CUDA moat in inference, requiring only eight keystrokes to move.

0
Q

How does the 750-megawatt data center deal with OpenAI help scale their cloud footprint?

Question 17 of 18
Andrew Feldman

OpenAI needed to secure data center capacity, which is measured in megawatts. We are delivering a full cloud solution, leasing two hundred fifty megawatts in 2026, 2027, and 2028. OpenAI connects directly to our clusters via an API to run their workloads.

0
Q

Why did TSMC agree to modify their manufacturing process for a 30-person startup in 2017?

Question 18 of 18
Andrew Feldman

We presented a bold proposal that allowed TSMC to leverage their strengths without changing too much. They recognized that AI would run better on giant chips and were willing to take a risk to learn. This collaborative spirit is how great companies win.

0

Table Of Questions

1Why has token speed become the dominant conversation in the AI sector?
2What does speed mean in terms of user experience metrics?
3How does the Netflix evolution analogy apply to the impact of fast inference?
4How do you define the different choices made in the specialized chip landscape?
5What did Nvidia's acquisition of Groq reveal about the GPU architecture's limits?
6What is Broadcom's Jalapeno chip, and how does it fit into OpenAI's strategy?
7What are the three major silicon manufacturing bottlenecks limiting GPU supply?
8Why does Cerebras's wafer-scale engine process AI workloads faster than standard GPUs?
9How does agentic AI drive an massive shortage of traditional CPUs?
10Why is building model architecture directly into silicon design a structural mistake?
11How did the team navigate the early years when the market was not ready?
12Why do standard GPUs struggle with data movement during the decode phase of inference?
13Why did Cerebras choose to build a dinner-plate-sized wafer-scale chip?
14What was the hardest packaging problem you had to solve during early chip building?
15How does Cerebras manage chip defects and yield reliability at the wafer scale?
16Why do you argue that Nvidia's CUDA software platform is no longer a durable moat?
17How does the 750-megawatt data center deal with OpenAI help scale their cloud footprint?
18Why did TSMC agree to modify their manufacturing process for a 30-person startup in 2017?

Video Interviews with Andrew Feldman

Cerebras CEO: CUDA Is Not a Moat | Andrew Feldman

Cerebras CEO: CUDA Is Not a Moat | Andrew Feldman

Cerebras CEO: CUDA Is Not a Moat | Andrew Feldman

Cerebras Systems IPO: Andrew Feldman on public debut

Cerebras Systems IPO: Andrew Feldman on public debut

Cite This Interview

Use this interview in your research, article, or academic work

Related Interviews

Travis Kalanick, Founder & CEO at Atoms
Technology

Travis Kalanick

Founder & CEO at Atoms

Not Publicly Disclosed
2000+

In this interview, Atoms founder and CEO Travis Kalanick shares his vision for automating the physical world through industrial robotics and AI. Kalanick reflects on…

Read Interview
David Baszucki, Founder & CEO at Roblox
Technology

David Baszucki

Founder & CEO at Roblox

Approx. 400 Million USD
3100+

In this interview, Roblox co-founder and CEO David Baszucki discusses the platform's evolution into a major digital human co-experience platform. Baszucki outlines the value of…

Read Interview
Garry Tan, President & CEO at Y Combinator
Technology

Garry Tan

President & CEO at Y Combinator

Not Publicly Disclosed
100+

In this interview, Y Combinator President and CEO Garry Tan outlines the new paradigms of software development driven by vibe coding and AI agents. Tan…

Read Interview
Ted Sarandos, Co-CEO at Netflix
Technology

Ted Sarandos

Co-CEO at Netflix

Approx. 3.75 Billion USD
16000

In this interview, Netflix Co-CEO Ted Sarandos discusses the platform's ten-year journey in India and its global expansion strategy. Sarandos explains how Netflix matches content…

Read Interview
Brian Schimpf, Co-founder & CEO at Anduril Industries
Technology

Brian Schimpf

Co-founder & CEO at Anduril Industries

Approx. $183.3 Million
7000+

In this interview, Anduril Industries co-founder and CEO Brian Schimpf discusses the shift in modern warfare toward AI-powered, software-first defense technologies. Schimpf shares the origin…

Read Interview
Amanda McMaster, Interim CEO at Boston Dynamics
Technology

Amanda McMaster

Interim CEO at Boston Dynamics

Not Publicly Disclosed
1500+

In this interview, Boston Dynamics interim CEO Amanda McMaster explains the commercial transition of the pioneer robotics firm from research laboratory to industrial deployer. McMaster…

Read Interview