**Please strictly adhere to the following resume naming convention:ALL CAPS, NO SPACES BETWEEN UNDERSCORESPTN_US_GBAMSREQID_CandidateBeelineIDExample: PTN_US_9999999_SKIPJOHNSON0413: - MSP Owner: Michelle LeeLocation: Woonsocket, RI / HybridDuration: 6 monthsGBaMS ReqID: 10945375Role SummaryWe are looking for a highly experienced Site Reliability Engineer to design, operate, automate, and continuously improve highly reliable, scalable, secure, and observable cloud GCP. The ideal candidate will bring deep hands-on experience in SRE practices, multi-cloud infrastructure, Kubernetes platforms, CI/CD engineering, infrastructure automation, observability, disaster recovery, incident response, and reliability governance. This role is suited for an engineer with strong experience in SLI/SLO management, error budget governance, cloud migration, platform modernization, production support, and automation-led toil reduction. The candidate will be responsible for improving platform availability, deployment reliability, incident response maturity, disaster recovery readiness, observability coverage, and operational efficiency across enterprise cloud environments. The role requires close collaboration with development, security, infrastructure, platform, and business teams to ensure business-critical applications meet defined reliability, performance, compliance, and scalability objectives.Key Responsibilities• Site Reliability Engineering and Reliability Governance (SLIs, SLOs, error budgets, reliability reviews, and blameless postmortem practices, MTTD, MTTR, recurring incidents)• Kubernetes, Containers, and Platform Engineering (Kubernetes, EKS, AKS, GKE, ECS, and Docker-based platforms)• Infrastructure as Code and Automation (Terraform)• CI/CD and Release Reliability (GitHub Actions, blue-green, canary, rolling deployments, automated rollback, deployment validation, and automated testing)• Observability, Monitoring, and Logging (Prometheus, Grafana)• Disaster Recovery, High Availability, and Resilience• Security, Compliance, and Cloud Governance• Linux Systems Administration and Production SupportRequired Qualifications• 10+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, Linux administration, production operations, or platform engineering.• Strong hands-on experience designing, operating, and supporting cloud infrastructure across AWS, Azure, and GCP.• Deep experience with Kubernetes platforms such as EKS, AKS, GKE, and containerization using Docker.• Experience building and maintaining CI/CD pipelines using Jenkins, GitHub Actions.• Strong observability experience with Prometheus, Grafana, ELK Stack, OpenSearch, Log Analytics, Application Insights, and GCP Cloud Monitoring.• Experience with disaster recovery, high availability, backup automation, multi-region failover, and recovery validation.• Hands-on scripting and automation experience using Python, Bash, PowerShell, and Ansible.• Linux systems administration experience across enterprise production environments.Preferred Qualifications• Experience implementing GitOps using ArgoCD, Helm.• Experience with service mesh and API traffic management using Istio Service Mesh and API Gateway.• Experience supporting regulated enterprise environments in healthcare, financial services, banking, insurance, or similarly controlled domains. The resume includes experience across financial, healthcare, insurance, retail, and banking clients.• Experience with cloud cost optimization.• Experience with security and governance tooling such as GuardDuty, CloudTrail, Kubernetes RBAC, Secrets Manager.• Experience authoring runbooks, DR playbooks, operational procedures, architecture diagrams, and incident response documentation.Certifications Preferred• Google Cloud Certified Cloud EngineerSkills: Core JavaExperience Required: 8-10
