Overview
What It Does
Gke Ai Troubleshooting Handle Disruption Gpu Tpu packages a focused devops & cloud workflow for an AI agent. Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE. Use when diagnosing node disruptions, predicting host maintenance events on GPU/TPU nodepools, inspecting node interruption PromQL metrics, auditing node taints, or configuring workload protection strategies (graceful termination, opportunistic maintenance, PodDisruptionBudgets). Don't use for general GKE cluster creation, network policy configuration, or non-disruption workload deployment. It is best suited to users who can review the resulting actions and provide only the accounts, files, or command access needed for the task. It is not a substitute for human approval on destructive, financial, security-sensitive, or public-facing actions.
Task ideas
Popular Use Cases
- Inspect deployment or infrastructure state
- Automate a repeatable operations task
- Troubleshoot configuration and runtime issues
- Apply documented production practices
Installation
Install this Agent Skill
Claude Code
npx skills add https://github.com/google/skills --skill gke-ai-troubleshooting-handle-disruption-gpu-tpuCommands derived from the public GitHub SKILL.md record. Checked 2026-09-02. Review the source before running them.
Before you start
Requirements
| Claude Code | Required / review |
| Codex | Required / review |
| Public SKILL.md source | Required / review |
| Review instructions and requested permissions before installation | Required / review |
| Paid service | Check source |
| Supported system | Check source |
Popularity context
Why It’s Popular
Gke Ai Troubleshooting Handle Disruption Gpu Tpu is a verified Agent Skill from google with a public SKILL.md, compatible with Claude Code, Codex, Gemini CLI.
Alternatives