Brazilian company hires for hybrid or remote position
Location: Brazil (any location)
- ️ Only candidates already based in Brazil will be considered
Work Model: Hybrid for candidates living in state capitals and Remote for candidates living in countryside/cities outside the state capitals
️ Language Requirements: Advanced/Fluent English – MandatoryMandatory (there will be direct contact with an international client)
Seniority: Senior (6+ years)
Compensation: Please inform your salary expectations when applying.
- ️ Instructions: Please send your CV in English and make sure to include all skills and experience that match the requirements of the opportunity. This will significantly increase your chances of success.
We are looking for an experienced Site Reliability Engineer (SRE) to improve the reliability, scalability, performance, and operational excellence of enterprise cloud environments.
You will work closely with development, platform, and infrastructure teams to automate operations, improve observability, optimize deployments, and ensure high availability across mission-critical systems.
If you enjoy solving complex infrastructure challenges through engineering, automation, and modern cloud technologies, this opportunity is for you.
We are seeking a highly skilled Site Reliability Engineer with extensive experience in cloud infrastructure, DevOps practices, automation, and distributed systems.
The ideal candidate combines strong infrastructure knowledge with software engineering principles, helping organizations improve operational efficiency through automation, Infrastructure as Code (IaC), observability, and continuous improvement.
You should be comfortable working in high-availability environments, responding to critical incidents, and continuously enhancing platform resilience.
You will be responsible for:
Automation & Operational Excellence
-
Automate operational processes, deployments, and infrastructure provisioning.
-
Develop internal automation tools using Python and Bash.
-
Design, implement, and optimize CI/CD pipelines.
-
Reduce repetitive operational activities through automation.
Cloud Infrastructure & Infrastructure as Code
-
Provision and manage cloud infrastructure.
-
Implement Infrastructure as Code using Terraform (preferred) and other IaC tools.
-
Ensure scalability, resilience, and high availability.
-
Design auto-scaling and self-healing solutions.
Reliability & Performance
-
Monitor applications and infrastructure using metrics, logs, and distributed tracing.
-
Define and monitor SLIs, SLOs, and SLAs.
-
Perform capacity planning.
-
Troubleshoot performance bottlenecks and system issues.
Incident Management
-
Respond to critical production incidents.
-
Conduct Root Cause Analysis (RCA).
-
Drive continuous improvement initiatives.
Observability
-
Implement monitoring, alerting, and observability platforms.
-
Build operational dashboards.
-
Improve end-to-end system visibility.
Platform Engineering
-
Collaborate closely with software engineering teams.
-
Support DevOps and SRE best practices.
-
Improve deployment processes, system architecture, and application resilience.
Business Continuity
-
Support Disaster Recovery strategies.
-
Execute load and stress testing.
-
Implement Chaos Engineering practices.
-
Improve operational resilience.
Governance
-
Define operational standards.
-
Document infrastructure and architecture.
-
Promote security best practices and DevSecOps principles.
- ️ All requirements below are mandatory and eliminatory. Candidates who cannot clearly demonstrate these qualifications in their CV are unlikely to proceed in the recruitment process.
Education
✔ Bachelor's Degree in:
-
Computer Science
-
Engineering
-
or a related field
-
Advanced/Fluent English.
-
Proven experience as a Site Reliability Engineer (SRE).
-
Strong background in:
-
Site Reliability Engineering
-
DevOps
-
Cloud Infrastructure
-
Cloud Engineering
-
Hands-on experience with at least one major cloud platform:
-
AWS
-
Microsoft Azure
-
Google Cloud Platform (GCP)
-
Practical experience with:
-
Terraform (preferred)
-
Docker
-
Kubernetes
-
CI/CD Pipelines
-
Python or Bash
-
Linux/Unix
-
Strong experience with monitoring and observability platforms.
-
Incident Management and Production Support.
-
Strong troubleshooting and problem-solving skills.
-
Experience building highly available and scalable systems.
The following qualifications will be considered a strong advantage:
-
Large-scale enterprise environments.
-
Consulting experience.
-
Chaos Engineering.
-
Highly distributed architectures.
-
Advanced observability platforms.
-
Apache Kafka.
-
Financial Services or mission-critical environments.
-
DevSecOps.
-
Platform Engineering.
-
Enterprise-scale cloud infrastructure projects.
-
Modern Site Reliability Engineering culture.
-
Cloud-native and Platform Engineering initiatives.
-
Strong DevOps and Infrastructure as Code practices.
-
Collaboration with international engineering teams.
-
High-impact role supporting mission-critical platforms.
-
Modern engineering environment focused on automation and operational excellence.
✅ Have I worked as a Site Reliability Engineer (SRE) or in an equivalent role focused on cloud reliability, automation, and operational excellence?
✅ Does my CV clearly demonstrate hands-on experience with AWS, Azure, or GCP, as well as Terraform, Docker, Kubernetes, Linux, Python/Bash, and CI/CD pipelines?
✅ Have I implemented observability solutions, managed production incidents, conducted Root Cause Analysis (RCA), and worked with SLIs, SLOs, and SLAs?
✅ Do I have experience designing highly available, scalable, and resilient cloud environments using Infrastructure as Code and DevOps best practices?
✅ Am I fluent in English and comfortable collaborating with international engineering teams in mission-critical production environments?
If you answered "No" to one or more of these questions, we recommend carefully reviewing your fit before applying.
This position is intended for a Senior Site Reliability Engineer (SRE) with extensive experience in cloud infrastructure, automation, observability, DevOps, and platform reliability.
Candidates whose experience is primarily focused on traditional infrastructure administration, system support, or operations without demonstrated expertise in Infrastructure as Code, Kubernetes, cloud platforms, automation, CI/CD, and Site Reliability Engineering practices are unlikely to meet the expectations for this role.
Site Reliability Engineer, SRE, DevOps Engineer, Platform Engineer, Cloud Engineer, Infrastructure Engineer, Cloud Infrastructure, AWS, Amazon Web Services, Microsoft Azure, Google Cloud Platform, GCP, Terraform, Infrastructure as Code, IaC, Docker, Kubernetes, Linux, Unix, Python, Bash, Shell Scripting, CI/CD, Jenkins, GitLab CI, GitHub Actions, Monitoring, Observability, Prometheus, Grafana, ELK Stack, OpenTelemetry, Distributed Tracing, Logging, Metrics, SLI, SLO, SLA, Incident Management, Root Cause Analysis, RCA, Capacity Planning, Auto Scaling, Self-Healing, Disaster Recovery, Chaos Engineering, DevSecOps, Platform Engineering, Apache Kafka, High Availability, Scalability, Enterprise Infrastructure, Financial Services
#EY BR