Co-founder & CEO at Cerebras Systems
AI inference is dominated by moving weights from memory to compute. By building a dinner-plate-sized chip stuffed with fast SRAM instead of HBM, we keep the weights directly on the silicon. This allows us to move data to compute two thousand five hundred times faster than a GPU.
This answer is part of a full interview with Andrew Feldman, Co-founder & CEO at Cerebras Systems.
Found this insight valuable? Share it with your network to help others learn from Andrew Feldman's experience.
Use this answer in your research, article, or academic work