Encora
April 2022 - present
Implemented observability solutions using Prometheus and Grafana, building dashboards and alerting rules based on SLOs, improving system visibility and reducing incident detection time. Designed and maintained CI/CD pipelines (Jenkins, GitHub Actions), improving deployment reliability and reducing failed deployments. Automated AWS infrastructure provisioning using Terraform, ensuring consistent, secure, and scalable environments across multiple services. Implemented security controls aligned with HIPAA/SOC2, including IAM policies, encryption, and audit logging, improving compliance posture and data protection. Improved monitoring and alerting strategies, reducing incident response time by ~30–40% and enhancing on-call efficiency. Led incident response processes, including root cause analysis and postmortems, driving continuous reliability improvements and preventing recurring issues. Defined and monitored SLOs/SLIs and introduced error budget practices to balance reliability and feature delivery. Developed automation scripts for patching, validation, and system health checks using Python, Bash, and Ansible, reducing manual operational effort. Conducted failover simulations and reliability testing, validating system resilience and improving disaster recovery readiness. Collaborated daily with U.S.-based teams in English across time zones, supporting production operations and incident resolution. Maintained runbooks, SOPs, and documentation, improving operational readiness and knowledge transfer. Supported cloud-native event-driven architectures using AWS services such as Kinesis, SQS, Lambda, and Step Functions for data processing and integration workflows. Implemented DevSecOps practices including IAM hardening, encryption, audit logging, and vulnerability-aware deployment processes.