
Big Data · MLOpsApplied academic project2026
Telco churn Lakehouse and MLOps
Reproducible Databricks Medallion Lakehouse and MLOps lifecycle with execution evidence.
Problem
Build a reproducible data and ML lifecycle for predicting customer churn while making the synthetic nature of the dataset explicit.
Architecture
Databricks Lakehouse with incremental ingestion, a Delta Lake Medallion architecture, temporal feature engineering, MLflow, Unity Catalog and monitored batch inference.
Data flow
Synthetic data → Auto Loader → Bronze/Silver/Gold → point-in-time dataset → training and registry → simulated inference and monitoring.
Results
16,316,445
Training rows
View source2,042,162
Distinct customers
View source33
Model features
View source2023-07 → 2024-12
Training window
View sourceContribution and authorship
- End-to-end co-implementation confirmed by Alonso on 2026-09-14
- MLflow experiment, three-task ML job and simulation run associated with Alonso
- Shared work across the Medallion pipeline, model lifecycle and daily validation
Alonso Marcos Muñoz · End-to-end co-authorJose Barros · Co-author
Evidence
Limitations
- • The data is synthetic and the production-labelled scenario is an academic simulation
- • Authenticated validation and deployment require an active Databricks workspace
- • Authorship is shared and both co-authors remain credited
Stack
DatabricksPySparkDelta LakeAuto LoaderDatabricks Asset BundlesMLflowUnity CatalogLakehouse Monitoring
https://github.com/AlonsoMarcosM/databricks-telco-churn-lakehouse
