Co-founder & CEO at Cerebras Systems
In graphics, you move data to the GPU, calculate for a long time, and send the result. Inference is the opposite: you move a massive volume of weights from memory to compute to calculate a single word, and then repeat. Traditional HBM memory bandwidth is too slow for this.
This answer is part of a full interview with Andrew Feldman, Co-founder & CEO at Cerebras Systems.
Found this insight valuable? Share it with your network to help others learn from Andrew Feldman's experience.
Use this answer in your research, article, or academic work