DigitalOcean
minorresolved

Serverless Inference - Gemma4 Latency Issues Causing Timeouts & Slow Responses

Jul 13, 2026, 07:04 PM → Jul 15, 2026, 09:46 PM · 2d 2h

Affected components

Inference

Update timeline

  1. resolvedJul 15, 2026, 09:46 PM

    The deployed fix has successfully restored full functionality, and our monitoring shows that system performance has completely stabilized. Response times for all Gemma 4 inference workflows have returned to normal baseline levels. We will continue to track platform stability moving forward to ensure long-term reliability. We apologize for any disruption this may have caused to your workflows and appreciate your patience throughout the recovery process.

  2. monitoringJul 15, 2026, 10:06 AM

    A fix has been deployed to resolve the issue. We are closely monitoring system performance to ensure full recovery and normal response times for all Gemma 4 inference workflows.

  3. identifiedJul 14, 2026, 11:07 AM

    We are currently experiencing an issue affecting customers using the Gemma 4 model on our Serverless Inference platform. Customers may experience significantly increased latency or request timeouts. Our Engineering team has identified a backend configuration issue as the root cause, which is temporarily impacting model performance. Please be assured that our Engineering team is actively working on a fix and is treating this issue with high priority. We sincerely apologise for any inconvenience

  4. monitoringJul 13, 2026, 10:24 PM

    A fix has been deployed to resolve the backend configuration issue. We are closely monitoring system performance to ensure full recovery and normal response times for all Gemma 4 inference workflows.

  5. identifiedJul 13, 2026, 07:04 PM

    We are currently experiencing an issue where customers using the Gemma 4 model on our Serverless Inference and Dedicated Inference platforms may experience severe latency or request timeouts. Our engineering team has identified a backend configuration issue as the root cause, which is temporarily degrading performance. We are actively working on a fix to restore normal response times and will provide another update as soon as the mitigation is in place