Simera Professional Key (SPK)

Marcell L

Peru

Data Engineer

$ 2,800/month

7 yrs exp

Skills

  • Azure
  • Docker
  • Git
  • GitHub
  • Java
  • MySQL
  • Python
  • Scala
  • Spark
  • SQL
  • AirFlow
  • Azure DevOps
  • BigQuery
  • Bitbucket
  • Cloud Functions
  • Cloud Run
  • Cloud Storage
  • Data Factory
  • DataBricks
  • DataFlow
  • Matplotlib
  • OpenCV
  • Pandas
  • PLSQL
  • Pyspark
  • Scikit-learn
  • Tensorflow
  • VertexAI
  • Python/R
  • Python3
  • FastAPI
  • pytorch
  • Generative AI

Marcell L

Peru

Data Engineer

$ 2,800 /month

7 yrs exp

Skills

  • Azure
  • Docker
  • Git
  • GitHub
  • Java
  • MySQL
  • Python
  • Scala
  • Spark
  • SQL
  • AirFlow
  • Azure DevOps
  • BigQuery
  • Bitbucket
  • Cloud Functions
  • Cloud Run
  • Cloud Storage
  • Data Factory
  • DataBricks
  • DataFlow
  • Matplotlib
  • OpenCV
  • Pandas
  • PLSQL
  • Pyspark
  • Scikit-learn
  • Tensorflow
  • VertexAI
  • Python/R
  • Python3
  • FastAPI
  • pytorch
  • Generative AI

Data Engineer

MINSAIT
January 2025 - December 2025

Migrated business logic from PL/SQL to PySpark improving scalability and performance of data processes. Automated deployment of PySpark notebooks using Jenkins CI/CD. Validated and reconciled historical data with data steward ensuring data integrity and consistency. Optimized PySpark processes and refactored code to improve efficiency and reduce technical debt by correcting bad practices.

Data Engineer

NTTDATA
November 2023 - December 2024

Developed Scala Jars for data ingestion, curation, and exploitation. Created process to extract CSV files from SFTP server to ADLS and load into Delta tables orchestrated in Databricks workflow. Configured Aercorsoft tasks for SAP table ingestion to ADLS in parquet format. Developed historical load process to Delta tables using liquid clustering. Merged CDC records with historical tables. Transformed data per business rules using PySpark. Orchestrated stored procedures and report exports with Airflow. Optimized BigQuery queries to reduce data retrieval time in PowerBI. Developed clustering models for customer churn using K Means and DBSCAN algorithms.

Data Engineer

Grupo Llyrod
February 2023 - September 2023

Ingested data from Oracle DB to BigQuery using Apache Beam and DataFlow. Migrated Oracle PL/SQL procedures to BigQuery. Standardized and homogenized data. Developed orchestration mesh for processes via Cloud Composer and Airflow. Developed ETL using DataFlow Apache Beam.

Machine Learning Engineer

Binareon S.A.C
June 2022 - September 2022

Implemented a vehicle license plate recognition system using computer vision and deep learning techniques.

Machine Learning Engineer

Instituto Geofísico del Perú
March 2020 - April 2021

Extracted, cleaned, analyzed, and visualized features from seismographs. Implemented algorithms like SMOTE and cost-sensitive methods to alleviate class imbalance. Classified seismic signals from Sabancaya volcano using ensemble of deep neural networks and random forests. Developed a FastAPI for batch predictions.

Research Assistant

Universidad Católica San Pablo
March 2017 - March 2019

Designed a regularization algorithm to mitigate overfitting in deep neural networks under data scarcity for regression and classification tasks.

Intern

Universidad Católica San Pablo
March 2016 - March 2017

Developed a mobile application presenting real-time data from a galvanic sensor with interactive dashboards to analyze human skin response to stimuli.

Smart Scores

Communication
70
Role Fit
95
Adaptability
80
Problem-solving
75

Smart Skills

beta