One faulty GPU worker failed to process requests

Partial outageResolvedVirtual Try-OnStarted Lasted 8 hours and 9 minutes

Updates

  1. Resolved

    • One GPU worker went into a bad state, it was restarted and returned to normal operation
    • A mechanism to detect and auto-restart such states was developed and deployed.
  2. Investigating

    Some of the requests to the nightly endpoint (experimental in app) are returning errors