BenchDuel logoBenchDuel
AI News

Cerebras Runs OpenAI Model at 750 Tokens Per Second

August 14, 2026
Cerebras Runs OpenAI Model at 750 Tokens Per Second

Cerebras just partnered with OpenAI to run flagship AI models at record speeds. Their special hardware hits up to 750 output tokens every single second. This makes the model up to fourteen times faster than standard processing.

This speed comes from a new option called the Ultrafast tier. It is currently available as a limited preview in the OpenAI API for select customers. More people will get access as they build more capacity.

This development shows how custom chips can beat traditional graphics cards in raw speed. Fast AI responses change how people build real-time apps. Users will no longer have to wait for long answers to appear.

Comments (0)

No comments yet. Be the first!

More AI news