google/skills

google-cloud-waf-operational-excellence

- Generates operations-focused guidance for Google Cloud workloads based on the design principles and recommendations in the Operational Excellence pillar of the Google Cloud Well-Architected Framework (WAF).

Voir la source
Document Skill original

Rendu depuis le dépôt source en conservant titres, exemples, code, tableaux, liens et images.

Google Cloud Well-Architected Framework skill for the Operational Excellence pillar

Overview

The operational excellence pillar in the Google Cloud Well-Architected Framework provides recommendations to operate workloads efficiently on Google Cloud. Operational excellence in the cloud involves designing, implementing, and managing cloud solutions that provide value, performance, security, and reliability. The recommendations in this pillar help you to continuously improve and adapt workloads to meet the dynamic and ever-evolving needs in the cloud.

Core principles

The recommendations in the operational excellence pillar of the Well-Architected Framework are aligned with the following core principles:

  • Ensure operational readiness: Define and measure criteria for a workload

to be considered ready for production, including staffing, processes, and governance. Grounding document: https://docs.cloud.google.com/architecture/framework/operational-excellence/operational-readiness-and-performance-using-cloudops.md.txt

  • Manage incidents and problems: Establish structured processes for

incident response, communication, and root cause analysis to minimize impact and prevent recurrence. Grounding document: https://docs.cloud.google.com/architecture/framework/operational-excellence/manage-incidents-and-problems.md.txt

  • Manage and optimize cloud resources: Monitor resource utilization and

right-size environments to maintain performance while ensuring operational efficiency. Grounding document: https://docs.cloud.google.com/architecture/framework/operational-excellence/manage-and-optimize-cloud-resources.md.txt

  • Automate and manage change: Use Infrastructure as Code (IaC) and CI/CD

pipelines to ensure consistent, repeatable, and low-risk deployments and configuration changes. Grounding document: https://docs.cloud.google.com/architecture/framework/operational-excellence/automate-and-manage-change.md.txt

  • Continuously improve and innovate: Regularly review architectures,

monitor industry trends, and adapt operations to meet evolving business needs. Grounding document: https://docs.cloud.google.com/architecture/framework/operational-excellence/continuously-improve-and-innovate.md.txt

Relevant Google Cloud products

The following are examples of Google Cloud products and features that are relevant to operational excellence:

  • Observability and monitoring
  • Cloud Monitoring: Full-stack observability for Google Cloud and

hybrid environments.

  • Cloud Logging: Real-time log management and analysis at scale.
  • Error Reporting: Aggregates and displays errors for running cloud

services.

  • Service Monitoring: Tools for defining and tracking Service Level

Objectives (SLOs).

  • Automation and CI/CD
  • Cloud Build: Serverless platform for building, testing, and

deploying software.

  • Cloud Deploy: Managed continuous delivery service for GKE, Cloud

Run, and GCE.

  • Terraform / Infrastructure Manager: Managed service for

Infrastructure as Code (IaC) automation.

  • Artifact Registry: Central repository for managing build artifacts

and container images.

  • Resource management and optimization
  • Recommender (Active Assist): Automatically identifies idle resources

and right-sizing opportunities.

  • Resource Manager: Hierarchical management of resources across

organizations, folders, and projects.

  • Incident response
  • Incident response & management (IRM): Structured tools and processes

for managing operational disruptions.

Workload assessment questions

Ask appropriate questions to understand operations-related requirements and constraints of the workload and the user's organization. Choose questions from the following list:

  • Operational readiness and performance
  • How do you define and measure operational readiness for your cloud

workloads and what specific criteria or metrics do you use?

  • Describe your process for defining, tracking, and achieving SLOs for

your critical workloads.

  • Incident and problem management
  • Describe your incident management process, including roles,

responsibilities, and communication channels.

  • How do you conduct post-incident reviews (PIRs) to identify root causes

and implement preventive measures?

  • Resource management and optimization
  • How do you ensure that your cloud resources are right-sized for your

workloads, and what tools or techniques do you use?

  • Change automation
  • Describe your change management process, including approval workflows,

testing procedures, and deployment strategies.

  • How do you automate deployments, ensure their consistency and manage

configuration?

  • Continuous improvement
  • How do you ensure that your cloud operations are continuously adapting

to meet evolving business needs and technological advancements?

Validation checklist

Use the following checklist to evaluate the architecture's alignment with operational excellence recommendations:

  • Operational readiness
  • [ ] A formal framework or set of criteria exists to assess operational

readiness before production deployment.

  • [ ] Service Level Objectives (SLOs) are explicitly defined and monitored

using automated tools.

  • Incident management
  • [ ] Incident response roles and communication channels are clearly

defined and documented.

  • [ ] A structured, blameless post-mortem process is followed for all

major incidents.

  • Change automation
  • [ ] All infrastructure changes are performed using Infrastructure as

Code (IaC) to ensure consistency.

  • [ ] CI/CD pipelines are integrated with automated testing for all

deployment changes.

  • Resource optimization
  • [ ] Resource utilization is regularly reviewed using recommendations

from Active Assist or performance data.

  • Culture of improvement
  • [ ] A documented strategy is in place for regularly reviewing and

adapting cloud operations to industry advancements.

du même dépôt

Autres Skills

Tous les Skills
google
Communauté

google-analytics-admin-api-basics

- Manages Google Analytics account and property settings, enables the Analytics Admin API via the Cloud CLI, lists accounts and properties, and manages data streams, custom dimensions, conversion events, and integrations. Use when you need to programmatically configure Google Analytics accounts, provision properties, manage data retention, configure Measurement Protocol secrets, or manage Firebase and Google Ads links.

installations
4
GitHub Stars
20,3 k
Mis à jour
22 sept.
google
Communauté

gke-workload-security

- Audits, configures, and hardens workload-level security controls for Google Kubernetes Engine (GKE) applications and namespaces. Covers running cluster security audits (auditcluster.sh), configuring Workload Identity Federation (impersonation, KSA/GSA binding, and pod setup), enforcing Network Policies (default-deny and Dataplane V2 logging), isolating high-risk pods inside GKE Sandbox (gVisor), enforcing Pod Security Standards (restricted labeling), and mounting Secret Manager secrets via CSI (SecretProviderClass). Use when auditing cluster security posture, isolating namespaces, applying pod security standards, setting up Workload Identity, or configuring network policies and secret volume mounts. Don't use for cluster-wide control plane security, RBAC hardening, Binary Authorization, Shielded Nodes, or enabling platform-level GKE add-ons (use gke-platform-security instead).

installations
3
GitHub Stars
20,3 k
Mis à jour
22 sept.
google
Communauté

google-ads-api-account-diagnostics

- Diagnoses Google Ads account performance issues such as conversion loss (value or volume), low lead flow/volume, and lost impression share (opportunities) due to ad rank, bids, or budgets. Use when troubleshooting sudden performance drops, analyzing campaign impression share metrics, investigating low lead flow, or searching for bidding and budget constraints. Don't use for setting up new campaigns, uploading conversion events directly, or general Google Mobile Ads SDK integration issues (use gma-android-integrate instead).

installations
4
GitHub Stars
20,3 k
Mis à jour
22 sept.
google
Communauté

google-ads-api-mcp-setup

Guides developers through downloading, configuring, and installing the official open-source Google Ads MCP Server. Use this skill when a user wants to connect their AI assistant (such as Gemini, Claude Code, or Cursor) to their Google Ads account to query campaigns or retrieve reporting metrics using natural language.

installations
4
GitHub Stars
20,3 k
Mis à jour
22 sept.