Data Engineer in financial data platforms

Gabriel Franco

I build and maintain financial data pipelines with SQL, Spark SQL, PySpark, Databricks, Delta Lake, and Python automation. My current work is focused on metadata-driven transformations, schema validation, legacy rule modernization, documentation, GitHub Actions CI validation, and reliable batch delivery.

Gabriel Franco, Data Engineer

Sao Paulo, Brazil. Open to remote data engineering opportunities.

Hands-on data engineering experience in banking environments.

I work close to business teams and backoffice operations, translating financial rules into maintainable SQL, Spark SQL, and PySpark processes. I care about reliability, schema validation, traceability, technical documentation, troubleshooting, and pipelines that other teams can trust.

2+years in data roles
Bankingfinancial and insurance data
LakehouseDatabricks and Delta Lake
Qualityschema validation and traceability

Tools I use to move data from raw inputs to business-ready outputs.

My strongest area is batch data engineering with SQL, Python, PySpark, Databricks, Delta Lake, metadata-driven processing, schema validation, and financial data rules.

Data Processing

SQL, PySpark, Python, Databricks notebooks, Delta Lake tables, transformations, joins, validations, and business rules.

SQL PySpark Python

Platforms

Azure Databricks, cloud lakehouse patterns, Oracle Exadata integration, on-premise to cloud modernization, and Git workflows.

Databricks Delta Lake Oracle

Operations

Batch orchestration concepts, job monitoring, logs, troubleshooting, GitHub Actions CI validation, controlled deploy flows, technical documentation, and business-facing delivery.

Batch Git GitHub Actions

From dashboards and automation to production data pipelines.

Data Engineer - BRQ Digital Solutions / F1rst-Santander

Financial MIS area. PySpark and Spark SQL pipelines on Databricks, metadata-driven SQL parameterization, schema validation, Delta Lake processing, legacy financial rule modernization, Oracle/Exadata integration, Python automation, GitHub Actions CI validation, controlled deploy flows, and technical documentation.

Data Analyst Intern - Santander Brasil

Insurance data platform. Databricks queries, SQL and PySpark transformations, Power BI dashboards, Python automation for files, data cleansing, normalization, and support for internal analytics demands.

Current work and learning roadmap.

I am rebuilding my public portfolio gradually. Today, the active project is a certification simulator. The next data engineering projects will be added only when they have real code, documentation, and technical decisions to show.

Active prototype

Certification Simulator App

A certification study app built with Next.js, TypeScript, Tailwind, and Supabase. The first focus is Databricks certification practice; the next step is adding analytics around attempts, weak topics, and learning progress.

Repository coming soon
Learning roadmap

Metadata-driven Financial Pipelines

Study path focused on turning complex financial rules into parameterized, testable, and documented Spark SQL/PySpark transformations.

Follow on GitHub
Learning roadmap

Spark Schema Validation

Study path focused on validating schemas, data types, write behavior, table compatibility, and parity between legacy logic and new pipelines.

Follow on GitHub

View my resume.

Choose the English version for international opportunities or the Portuguese version for Brazilian recruiters and local processes.

Let us talk about data engineering.

I am interested in remote opportunities, data engineering projects, and teams building reliable cloud data platforms.