Simera Professional Key (SPK)

Rafael D

Brasil

Site Reliability Engineer

$ 6,200/month

10+ yrs exp

Professional with extensive IT experience in SRE, Infrastructure, Cloud, DevOps, and Platform Engineering. Expertise in the administration and evolution of high-availability environments, with expertise in AWS, GCP, Kubernetes, Terraform, Kafka, CI/CD, observability, and automation. Experience with distributed platforms, microservices, messaging, Linux/Windows servers, and critical environments, with a focus on reliability, scalability, performance, and security.

Skills

  • Amazon S3
  • AWS
  • Datadog
  • DevOps
  • Grafana
  • Kubernetes
  • New Relic
  • Prometheus
  • Sas Base
  • Terraform
  • Agile Methodology
  • Amazon SQS
  • GCP
  • Rabbit
  • Reliability Engineering
  • Agile Methodologies
  • Agile/Scrum Methodologies
  • SRE
  • RabbitMQ
  • Google Cloud Platform (GCP)
  • Agile methodologies (SCRUM, Kanban)
  • Agile Methodologies
  • Google Cloud Platform, Azure
  • Site Reliability Engineering (SRE)
  • Amazon Web Services (AWS)
  • Methodologies
  • KPI Monitoring
  • GitHub Actions
  • Amazon Q
  • Apache Kafka
  • CI/CD
  • kubernetes
  • Observability
  • Confluent Kafka

Rafael D

Brasil

Site Reliability Engineer

$ 6,200 /month

10+ yrs exp

Professional with extensive IT experience in SRE, Infrastructure, Cloud, DevOps, and Platform Engineering. Expertise in the administration and evolution of high-availability environments, with expertise in AWS, GCP, Kubernetes, Terraform, Kafka, CI/CD, observability, and automation. Experience with distributed platforms, microservices, messaging, Linux/Windows servers, and critical environments, with a focus on reliability, scalability, performance, and security.

Skills

  • Amazon S3
  • AWS
  • Datadog
  • DevOps
  • Grafana
  • Kubernetes
  • New Relic
  • Prometheus
  • Sas Base
  • Terraform
  • Agile Methodology
  • Amazon SQS
  • GCP
  • Rabbit
  • Reliability Engineering
  • Agile Methodologies
  • Agile/Scrum Methodologies
  • SRE
  • RabbitMQ
  • Google Cloud Platform (GCP)
  • Agile methodologies (SCRUM, Kanban)
  • Agile Methodologies
  • Google Cloud Platform, Azure
  • Site Reliability Engineering (SRE)
  • Amazon Web Services (AWS)
  • Methodologies
  • KPI Monitoring
  • GitHub Actions
  • Amazon Q
  • Apache Kafka
  • CI/CD
  • kubernetes
  • Observability
  • Confluent Kafka

Site Reliability Engineer

CARDOSO CLOUD & SRE LTDA
June 2026 - August 2026

Providing high-level technology consulting and engineering services through my own firm, CARDOSO CLOUD & SRE LTDA. Focused on delivering production-grade cloud solutions, infrastructure automation, and robust system resilience for international enterprise clients. Key Responsibilities & Deliverables: • Cloud Infrastructure Architecture: Designing, building, and optimizing scalable, highly available architectures on AWS and GCP using infrastructure as Code (IaC) with Terraform. • Containerization & Orchestration: Managing and maintaining high-performance Kubernetes clusters (EKS/GCP) and Docker environments to ensure seamless deployment and scaling. • Observability & Monitoring: Implementing advanced monitoring, logging, and alerting systems using Prometheus and Grafana to track SLA/SLO compliance and accelerate root-cause analysis. • CI/CD Pipelines: Engineering, automation, and optimizing continuous integration and delivery pipelines (Jenkins, Docker) to guarantee safe, fast, and repeatable deployments. • System Resilience & Reliability: Proactively monitoring system health, managing incident response strategies, and optimizing system architecture to prevent downtime and guarantee reliability.

Senior Site Reliability Engineer

Stone
July 2024 - March 2026

Responsible for managing, maintaining, and improving the corporate messaging platform, working with Apache Kafka (Confluent Platform, Confluent Cloud, and Amazon MSK), Google Pub/Sub, Amazon SNS, and other distributed messaging solutions in high-availability environments; Working across the entire lifecycle of messaging services, including provisioning, configuration, monitoring, troubleshooting, performance management, and operational governance; Providing specialized technical support to development and architecture teams, working in a cross-functional model to ensure scalability, resilience, and efficient integration between systems; Implementing and maintaining observability and monitoring using Datadog, New Relic, and native infrastructure tools, ensuring visibility and fast incident detection; Managing Kubernetes environments, supporting the deployment, operation, and scaling of applications and messaging components in cloud-native architectures; Developing and maintaining Infrastructure as Code (IaC) with Terraform, promoting automation, standardization, traceability, and governance of tech environments; Leading technical initiatives for messaging standardization and automation, driving the migration of services like SQS and RabbitMQ to the internal Karavela platform (Backstage), centralizing management and simplifying resource provisioning. Key Achievements: ✅ Led the standardization and automation of messaging services, reducing operational complexity and speeding up resource provisioning for development teams. ✅ Optimized Kafka infrastructure with Terraform automation and right-sizing strategies, cutting operational costs and boosting resource efficiency.

Site Reliability Engineer

Ebury Bank
February 2023 - March 2024

Working as a Site Reliability Engineer in the Accounts Squad, ensuring high availability, reliability, performance, and scalability for critical financial applications and services; Taking an active role in the company's tech and operational transition during the migration from Bexs to Ebury, supporting system integration, environments, and business processes; Managing and supporting production and non-production infrastructure, ensuring operational stability and adherence to reliability best practices; Implementing and maintaining monitoring, observability, and alerting using SRE tools for continuous service health tracking and proactive incident detection; Supporting development teams in implementing DevOps practices, Continuous Integration (CI), Continuous Delivery (CD), and operational process automation. Key Achievements: ✅ Refined the monitoring strategy, eliminating false positives and unnecessary alerts, which improved notification accuracy and reduced alert fatigue. ✅ Contributed to the transition and integration of environments during Bexs' acquisition by Ebury, ensuring business continuity for key services.

Senior Site Reliability Engineer

Tembici
March 2021 - January 2023

Managing and expanding cloud environments on AWS and Google Cloud Platform (GCP), ensuring high availability, security, and performance for enterprise applications; Implementing and maintaining Infrastructure as Code (IaC) using Terraform, promoting automation, standardization, versioning, and environment governance; Orchestrating microservices on Kubernetes, working on configuration, scalability, observability, and workload optimization in production; Developing Continuous Integration and Continuous Delivery (CI/CD) pipelines using Git, GitHub Actions, and DevOps practices; Integrating quality and security tools into the development cycle, using SonarCloud for code quality and coverage analysis, Dependabot for vulnerability management, and OWASP ZAP (ZapScan) for application security testing; Working in cross-functional, multidisciplinary teams, collaborating with development, product, security, and infrastructure teams to continuously evolve platforms. Key Achievements: ✅ Implemented standardized pipelines with quality checks, automated testing, and security tools (SAST/DAST), strengthening software development governance. ✅ Led container right-sizing initiatives and HPA optimization, boosting operational efficiency and reducing infrastructure costs.

Site Reliability Engineer

TOTVS
December 2019 - March 2021

Working with the Fluig product, contributing to the reliability, availability, and performance of enterprise applications in mission-critical environments; Starting in a dedicated SRE team and later joining the development teams, promoting DevOps practices and strengthening collaboration between operations and software engineering; Managing and supporting AWS cloud environments, ensuring scalability, stability, and operational efficiency for all services; Managing and orchestrating microservices using Kubernetes, handling configuration, monitoring, troubleshooting, and workload optimization in production; Implementing and maintaining Continuous Integration and Continuous Delivery (CI/CD) pipelines using Bamboo and Git, automating build, testing, and deployment processes. Key Achievements: ✅ Led infrastructure planning and sizing for peak demand periods, including critical Black Friday operations, ensuring full service availability. ✅ Implemented FinOps practices and right-sizing strategies based on monitoring metrics, optimizing cloud resource usage and cutting operational costs.

SysAdmin

TOTVS
October 2010 - December 2019

Managing and supporting highly critical enterprise monolithic applications, ensuring performance and stability across production, staging, and development environments; Managing application servers like JBoss, Apache Tomcat, and WildFly, handling installation, configuration, updates, troubleshooting, and performance tuning; Administering Oracle, Microsoft SQL Server, and Progress databases, working on monitoring, maintenance, support, and performance analysis; Managing Windows Server and Linux servers on-premises, including provisioning, configuration, patching, monitoring, and resource management; Developing and implementing automation for routine operations, reducing manual tasks and boosting infrastructure process efficiency; Implementing, managing, and improving Continuous Integration and Continuous Delivery (CI/CD) pipelines using Jenkins, driving faster and higher-quality software delivery. Key Achievements: ✅ Modernized and optimized CI/CD pipelines, reducing build times by roughly 75%—from over an hour down to about 15 minutes. ✅ Scaled up operational and deployment automation, boosting team productivity and cutting manual errors during releases.

Smart Scores

Communication
85
Role Fit
95
Adaptability
90
Problem-solving
90

Smart Skills

beta