Data & Machine Learning Engineering

José Paz Rangel Rojas

I build the data systems a real estate and media group in Mexico runs on: incremental ELT into Snowflake, anti-money-laundering regulatory reporting, geospatial dashboards and unsupervised anomaly detection. Actuary by training, which is mostly why I care whether a number can be defended.

8Production systems
R · PythonBoth, in production
SnowflakeWarehouse & Cortex
LFPIORPIAML compliance domain

Projects

Every demo runs on synthetic data and needs no credentials. Free instances sleep — the first load takes about 40 seconds.

Property Portfolio — Geospatial

Live demo

A 300-property portfolio on one map. Surface area is measured geodesically from each parcel's own polygon instead of trusted from a field that was often missing or wrong, and acquisition prices are restated for inflation so a portfolio total means something. Built twice — R/Shiny and Python/Dash — over the same warehouse, plus a natural-language query box on a Snowflake Cortex agent.

RShiny PythonDash Leafletpyproj Snowflake Cortex

Land Valuation Model

Live demo

What is this parcel worth, and how sure? A valuation model over 2,600 land transactions that returns a calibrated range, not a number — 90% promised, 90.2% measured. Prices are restated for inflation first, because otherwise the model learns that recent is expensive and calls it value. Pointed at the group's own portfolio, it finds the purchases made outside the market.

Pythonscikit-learn RegressionConformal prediction Streamlit

Retail Space Manager

Live demo

Leasing system for three shopping centres: floor plan, availability, tenants and history. Handing a unit back never deletes anything — it posts a negative-square-metre entry, so occupancy is the running sum and the history cannot drift from the current state. Drag the date slider and the plan redraws for any month since 2022, with no historical tables.

PythonStreamlit Append-only ledgerAltair Domain modeling

AML Anomaly Detection

Live demo

Nobody ever labelled a transaction as laundering, so a classifier is off the table. This learns what is normal for each client and ranks what departs from it. contamination comes from the team's actual review capacity, validation is an expanding window, and the evaluation is reported twice — with labels, and the way production will have to, without them.

Pythonscikit-learn Isolation ForestStreamlit Unsupervised

ERP → Snowflake ELT

Runnable locally

24 entities pulled from a paginated ERP API into Snowflake from one declarative catalog, merged by row hash so a re-stamped modification date writes nothing. Includes a fingerprint reconciler, data-quality views and drift probes. The demo runs four passes against SQLite: pass three proves the merge is idempotent.

PythonSnowflake ELTIdempotent merge Data quality

AML Regulatory Connector

Live demo

End-to-end reporting for Mexico's LFPIORPI: ERP to Snowflake to a Streamlit app behind AWS Cognito. The hard part was not the API — it was modelling partial payments, VAT and multi-invoice collections so the reported amount is the one that actually changed hands, instead of over-reporting and having to amend.

PythonStreamlit SnowflakeAWS Cognito DockerRegTech

Bank Reconciliation — Host-to-Host

Runnable locally

An hourly bank feed reconciled against ERP invoices. Runs on a six-hour overlapping window because banks backfill, and stays idempotent so the overlap costs nothing. Extracts tax IDs out of free-text bank descriptions and matches instalments and grouped payments.

RShiny SnowflakeBanking API Reconciliation

Bank Statement Consolidator

Runnable locally

My first project here, and the one that started the rest: an Excel VBA macro rebuilt as an R/Shiny app. Consolidates statements from multiple accounts whose headers never sit on the same row, and classifies every movement through 78 rules. 525 movements, none left unclassified.

RShiny VBA migrationETL

About these repositories

Six of the eight are real production systems, rewritten for public release. Companies, tenants, properties, tax IDs, bank accounts, people and vendor endpoints are invented; credentials live in environment variables and are not included; sample data is generated and matches the shape of the real data, not its content. The architecture, the business rules and the engineering decisions are the real ones — that is the part worth showing. Retail Space Manager is the one rebuild: same data model and same rules, ported from R/Shiny to Python, with three defects of the original fixed and documented.

The other two — AML Anomaly Detection and Land Valuation Model — are different, and each says so at the top of its README: they are reference implementations on generated data. The modeling decisions are the ones I would deploy; the deployments do not exist yet.