Simera Professional Key (SPK)

Sebastian B

Peru

Sr. Data Engineer

$ 3,700/month

6 yrs exp

Senior Data Engineer with 5+ years of experience designing and maintaining scalable ETL/ELT pipelines, data warehouses, data lakes, and lakehouse architectures in high-demand cloud environments. Advanced proficiency in SQL, Python, and PySpark for dimensional modeling, large-scale data transformation, and performance optimization. Proven experience with AWS (Glue, Lambda, S3, Redshift), Azure (Data Factory, Databricks, Synapse, Event Hubs), and Snowflake, orchestrating workflows with Apache Airf…

Skills

  • Amazon S3
  • AWS
  • Azure
  • CI/CD
  • data warehouse
  • Git
  • Python
  • Spark
  • SQL
  • Amazon AWS
  • Amazon Redshift
  • Amazon SQS
  • AWS Lambdas
  • Azure DevOps
  • Data Engineer
  • Data Warehousing
  • Phyton
  • Power BI
  • Pyspark
  • Snowflake
  • Python/R
  • CICD
  • Data engineering
  • PowerBI
  • Python3
  • Paython
  • debt
  • RAG
  • Langchain
  • dbt
  • ETL/ELT pipelines
  • MSQL
  • ETL/ELT Pipeline
  • Amazon Q
  • Data Quality
  • Apache Airflow
  • SOQL
  • Observability
  • Azure data factory
  • Amazon ECS
  • Datawarehouse
  • Azure Databricks
  • AWS Lambda
  • Azure Synapse

Sebastian B

Peru

Sr. Data Engineer

$ 3,700 /month

Part Time: $ 2,450/month

6 yrs exp

Senior Data Engineer with 5+ years of experience designing and maintaining scalable ETL/ELT pipelines, data warehouses, data lakes, and lakehouse architectures in high-demand cloud environments. Advanced proficiency in SQL, Python, and PySpark for dimensional modeling, large-scale data transformation, and performance optimization. Proven experience with AWS (Glue, Lambda, S3, Redshift), Azure (Data Factory, Databricks, Synapse, Event Hubs), and Snowflake, orchestrating workflows with Apache Airf…

Skills

  • Amazon S3
  • AWS
  • Azure
  • CI/CD
  • data warehouse
  • Git
  • Python
  • Spark
  • SQL
  • Amazon AWS
  • Amazon Redshift
  • Amazon SQS
  • AWS Lambdas
  • Azure DevOps
  • Data Engineer
  • Data Warehousing
  • Phyton
  • Power BI
  • Pyspark
  • Snowflake
  • Python/R
  • CICD
  • Data engineering
  • PowerBI
  • Python3
  • Paython
  • debt
  • RAG
  • Langchain
  • dbt
  • ETL/ELT pipelines
  • MSQL
  • ETL/ELT Pipeline
  • Amazon Q
  • Data Quality
  • Apache Airflow
  • SOQL
  • Observability
  • Azure data factory
  • Amazon ECS
  • Datawarehouse
  • Azure Databricks
  • AWS Lambda
  • Azure Synapse

Sr. Data Engineer

Cheers Health
July 2025 - May 2026

Designed and developed scalable ETL/ELT pipelines integrating Amazon, Walmart, and TikTok APIs into AWS analytics environments, processing 4M+ records daily with data-quality validation and workflow documentation. Implemented a lakehouse architecture on AWS (Lambda, Glue, S3, Redshift) with structured transformation layers plus production pipeline monitoring and observability. Led the migration of legacy pipelines (Xplenty.io) to AWS, modernizing the ingestion architecture and reducing overnight processing windows from 6+ hours to under 45 minutes. Implemented Type 2 SCD patterns for churn analysis and data-processing pipelines for AI models (LangChain, FAISS, OpenAI GPT).

Sr. AWS Data Developer

The Hawkers Club
January 2024 - May 2025

Built and optimized Snowflake analytics layers with dbt and Apache Airflow (100+ DAGs), consolidating CRM and operational sources into dimensional models that process 800K+ daily transactions. Designed data models and KPI dashboards in Power BI, translating business requirements for churn, LTV, and membership economics into repeatable metrics for non-technical stakeholders. Implemented data quality and integrity controls plus observability checks, reducing pipeline incidents from about three per day to zero and ensuring data availability and reliability. Automated document-classification workflows using LangChain, RAG, and OpenAI GPT APIs, reducing manual validation from about eight hours to under one hour per day.

Sr. Implementation Consultant

DigitalCALA
April 2021 - January 2024

Architected and implemented a medallion data platform (Bronze/Silver/Gold) in Azure Databricks and PySpark, processing 10M+ Telcel subscriber records with Azure Data Factory and Azure Event Hubs. Reduced data latency from 24 hours to under two hours and optimized Spark jobs from three hours to under 40 minutes through pipeline redesign and partition tuning. Applied Type 2 SCD patterns, idempotent pipelines, and data-governance standards; implemented CI/CD through Azure DevOps Pipelines and version control with Git. Mentored junior engineers in scalable design patterns, PEP 8 standards, and Git-based development workflows.

Data Analyst

Pallon
October 2020 - April 2021

Developed ETL pipelines in Python and SQL to feed computer-vision models using 50K+ infrastructure inspection images, reducing preprocessing time from 6+ hours to under one hour. Built analytical datasets and anomaly-detection pipelines for infrastructure-damage detection workflows.

Smart Scores

Communication
68
Role Fit
95
Adaptability
68
Problem-solving
64
Professional Presence
90
Drive/Initiative
90

Smart Skills

beta