On August 27 between 18:10 and 18:22 UTC, audio transcription processing was delayed for our customers leveraging our pooled inference infrastructure.
Jobs submitted during this window were queued and processed automatically once capacity was restored — no data was lost, no job was rejected, and no action is required from our customers and users.
The delay occurred during a routine scale-up operation when our cloud GPU provider had a temporary capacity shortage.
We have restored full capacity and are widening the range of hardware we can fall back to so this cannot recur.
Resolved
On August 27 between 18:10 and 18:22 UTC, audio transcription processing was delayed for our customers leveraging our pooled inference infrastructure.
Jobs submitted during this window were queued and processed automatically once capacity was restored — no data was lost, no job was rejected, and no action is required from our customers and users.
The delay occurred during a routine scale-up operation when our cloud GPU provider had a temporary capacity shortage.
We have restored full capacity and are widening the range of hardware we can fall back to so this cannot recur.
Monitoring
We've widened the list of eligible hardware, which resulted in immediate allocation with our Cloud Provider, and services are now back up. We're monitoring the job queue and confirming historical jobs enqueued during the incident are now properly processed.
Identified
We've identified a GPU shortage in our preferred region as the root cause of the incident. Jobs are still accepted and properly queued. They will be resolved with a slight delay.
Investigating
Audio transcription processing is currently delayed, as our GPU fleet failed to scale up with demand. This affect our customers leveraging our pooled inference infrastructure only.