Caro usuário, habilite o javascript para que esse site funcione corretamente.

Site Reliability Engineer (SRE) – GCP Platform

Pessoa JurídicaPresencial (Local)VIPSão Paulo-SPEmpresa Confidencial (Cadastre-se)

* Salário: R$ 2.000 a R$ 5.000 por mês (estimado)

* O valor exibido é uma estimativa calculada com base em dados públicos e referências do mercado. Não garantimos que este seja o salário oferecido para esta vaga específica.

Área: Outros

Nível: Pleno

About the job

Summary

Under general supervision, the Site Reliability Engineer (SRE) is responsible for improving the reliability, scalability, and resilience of digital platforms and infrastructure. Operating at the intersection of software engineering and systems engineering, this role focuses on building automation to eliminate toil (repetitive, manual tasks) and proactively preventing service-impacting incidents.

The SRE will design, build, and support large-scale, distributed, fault-tolerant systems in Google Cloud Platform (GCP). This role ensures that critical business platforms maintain high availability and performance while supporting a rapid pace of feature deployment. By championing observability and taking a holistic view of system health, the SRE will enhance cloud-based transformation initiatives, ensuring technology capabilities remain ahead of evolving customer needs and business growth.

Job Duties

  • Automation & Toil Reduction: Design, build, and maintain automation tools and infrastructure-as-code (IaC) to eliminate manual operational tasks and streamline deployment workflows.
  • Service Level Management: Collaborate with product and development teams to define, monitor, and enforce Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for critical services.
  • Observability & Monitoring: Design and implement robust monitoring, alerting, and logging systems to ensure deep visibility into application and infrastructure health.
  • Incident Management & Post-Mortems: Participate in an on-call rotation to provide operational support. Lead and participate in blameless post-mortems to identify root causes and implement preventative measures.
  • Cloud Architecture & Scalability: Optimize GCP cloud infrastructure for performance, cost, and scalability. Manage load balancing, traffic routing, and auto-scaling configurations.
  • CI/CD Pipeline Engineering: Support and enhance continuous integration and continuous deployment (CI/CD) pipelines to ensure safe, predictable, and frequent software releases.
  • Disaster Recovery & Resilience: Plan, implement, and test disaster recovery, failover, and backup strategies to guarantee business continuity.
  • Collaboration & Code Review: Review application architecture and code to ensure adherence to production readiness standards, security best practices, and scalability guidelines.

Education & Experience

  • Typically requires a bachelor’s degree in computer science, Information Technology, or a related field, 5 to 8 years of related experience in an SRE, DevOps, or Cloud Systems Engineering role; or an equivalent combination of education and experience.

Knowledge, Skills, Abilities

Technical Skills (SRE & GCP Core):

  • Cloud Platform: Experience deploying and managing workloads in Google Cloud Platform (GCP). Familiarity with core services such as GKE (Google Kubernetes Engine), Compute Engine, Cloud Run, Cloud Functions, and Big Query.
  • Containers & Orchestration: Strong understanding of Docker and Kubernetes management, deployment, and troubleshooting.
  • Infrastructure as Code (IaC): Hands-on experience with tools like Terraform to provision and manage cloud infrastructure.
  • Linux/Unix Systems: Deep knowledge of Linux/Unix administration, networking protocols (TCP/IP, DNS, HTTP/S), and system internals.
  • Programming/Scripting: Proficiency in at least one scripting or programming language used for automation (e.g., Python, Go, or Bash).
  • Observability Stack: Experience with monitoring and logging tools, specifically GCP Cloud Monitoring/Logging (Stack driver), or open-source equivalents like Prometheus, Grafana, and the ELK stack.
  • CI/CD Tools: Experience with modern deployment pipelines and version control tools (e.g., GitHub Actions, GitLab CI, Jenkins, or Google Cloud Build).

Soft Skills & Process:

  • Problem-Solving: Excellent troubleshooting and analytical skills, with a proven ability to debug complex, distributed systems under pressure.
  • Blameless Culture: Understanding of SRE culture, focusing on systemic fixes rather than human error during incident retrospectives.
  • Change Management: Demonstrated adherence to modern Change Management and GitOps workflows to ensure safe production changes.
  • Communication: Ability to interface well and collaborate effectively with software developers, product managers, security teams, and business stakeholders.

Physical Demands / Certifications

  • Preferred Certifications: Google Cloud Certified Professional Cloud DevOps Engineer, Professional Site Reliability Engineer, or Certified Kubernetes Administrator (CKA) is a strong plus.