CI&T
Brazil
[Job-32122] Site Reliability Engineer
RemotoHomeoffice
📍 BrazilRemotoPublicada 05 de outubro de 2026 · recém publicadacapturada em 05/10/2026
ANÁLISE DE COMPATIBILIDADE COM IA
Será que seu currículo passa pelo filtro desta vaga?
Envie seu CV e descubra, em segundos, sua compatibilidade real com [Job-32122] Site Reliability Engineer — o que já atende, e onde seu currículo pode estar perdendo aderência.
100% grátis
Resultado em segundos
Sem criar conta pra ver o preview
grátis · leva menos de 1 minuto
Descrição da vaga
At CI&T, we help large enterprises transform the potential of AI into real business impact with AI Deployment, AI-native execution, and tech-integrated business solutions.
With 30 years of experience in technological transformation, we accelerate innovation with expertise in Agentic SDLC, Application modernization, Data & AI, Martech and Business strategy.
We are 8,000 CI&Ters across more than 25 countries, collaborating to build solutions with real impact. AI is already part of how we work, evolve, and innovate every day.
Responsibilities
• Own the day-to-day operation of a monitoring platform, including dashboards, monitors, log pipelines, APM instrumentation, synthetic tests, and Real User Monitoring, keeping coverage accurate and current as our services evolve.
• Proactively sift through logs, error tracking, traces, and metrics to find failures, regressions, and anomalies that have not triggered an alert, then triage, reproduce, and drive them to resolution with the owning team.
• Tune alert thresholds, monitor logic, and notification routing to reduce noise and false positives while ensuring genuine customer-impacting issues page the right people quickly.
• Instrument new and existing services with meaningful metrics, structured logs, and distributed traces, and partner with engineers to improve the observability of their code.
• Write post-incident reviews, track remediation items to completion, and feed lessons learned back into monitors, runbooks, and system design.
• Define, measure, and report on service level indicators for key customer-facing services.
• Improve the reliability, scalability, and cost efficiency of our cloud infrastructure, CI/CD pipelines, and release processes, automating repetitive operational work wherever possible.
• Maintain and improve runbooks, escalation paths, and operational documentation so that the team can act quickly and consistently.
• Partner with our enterprise InfoSec team to remediate cybersecurity risk items across our infrastructure, including SSL/TLS cleanup, removal of exposed technology and version banners, implementation of security headers (HSTS, CSP, X-Frame-Options, and related), and DNS configuration hygiene (DNSSEC, SPF/DKIM/DMARC, dangling records), tracking findings from scans and audits through to verified closure.
• Collaborate with engineering, QA, product, and security teams to build reliability and observability into new features before they ship.
• Proactively sift through logs, error tracking, traces, and metrics to find failures, regressions, and anomalies that have not triggered an alert, then triage, reproduce, and drive them to resolution with the owning team.
• Tune alert thresholds, monitor logic, and notification routing to reduce noise and false positives while ensuring genuine customer-impacting issues page the right people quickly.
• Instrument new and existing services with meaningful metrics, structured logs, and distributed traces, and partner with engineers to improve the observability of their code.
• Write post-incident reviews, track remediation items to completion, and feed lessons learned back into monitors, runbooks, and system design.
• Define, measure, and report on service level indicators for key customer-facing services.
• Improve the reliability, scalability, and cost efficiency of our cloud infrastructure, CI/CD pipelines, and release processes, automating repetitive operational work wherever possible.
• Maintain and improve runbooks, escalation paths, and operational documentation so that the team can act quickly and consistently.
• Partner with our enterprise InfoSec team to remediate cybersecurity risk items across our infrastructure, including SSL/TLS cleanup, removal of exposed technology and version banners, implementation of security headers (HSTS, CSP, X-Frame-Options, and related), and DNS configuration hygiene (DNSSEC, SPF/DKIM/DMARC, dangling records), tracking findings from scans and audits through to verified closure.
• Collaborate with engineering, QA, product, and security teams to build reliability and observability into new features before they ship.
What We're Looking For:
• 3+ years of experience in site reliability engineering, DevOps, platform engineering, or a production-focused software engineering role.
• Hands-on experience administering and building in platform such as New Relic, Grafana, Splunk, or Dynatrace including dashboards, monitors, log management, and APM.
• Strong troubleshooting and root-cause analysis skills, with a demonstrated ability to work through logs, traces, and metrics to find the real problem behind a symptom.
• Working proficiency in at least one scripting or programming language such as Python, TypeScript/JavaScript, Go, or Bash, and comfort reading application code to understand failures.
• Experience operating services in a major cloud provider (AWS, GCP, or Azure), with a solid grasp of networking, containers, and Linux fundamentals.
• Familiarity with infrastructure as code (Terraform, CloudFormation, or Pulumi) and CI/CD tooling such as GitHub Actions.
• Experience with on-call responsibilities, incident response, and post-incident review processes.
• Clear written and verbal communication, including the ability to explain reliability concerns and trade-offs to non-technical stakeholders.
• Hands-on experience administering and building in platform such as New Relic, Grafana, Splunk, or Dynatrace including dashboards, monitors, log management, and APM.
• Strong troubleshooting and root-cause analysis skills, with a demonstrated ability to work through logs, traces, and metrics to find the real problem behind a symptom.
• Working proficiency in at least one scripting or programming language such as Python, TypeScript/JavaScript, Go, or Bash, and comfort reading application code to understand failures.
• Experience operating services in a major cloud provider (AWS, GCP, or Azure), with a solid grasp of networking, containers, and Linux fundamentals.
• Familiarity with infrastructure as code (Terraform, CloudFormation, or Pulumi) and CI/CD tooling such as GitHub Actions.
• Experience with on-call responsibilities, incident response, and post-incident review processes.
• Clear written and verbal communication, including the ability to explain reliability concerns and trade-offs to non-technical stakeholders.
Curtiu a vaga? Veja se seu CV se encaixa.
Sobe seu currículo lá em cima e receba seu score em segundos.
ALERTA DE VAGA CERTA
Quer receber vagas como essa antes da multidão do LinkedIn? De graça.