# Spark, Kafka and Airflow data platform

> Reproducible local platform with batch, streaming, Medallion architecture and orchestration.

- Category: Data Engineering
- Año / Year: 2026

## Context

Maintain a living metadata catalogue by combining structured, semi-structured and real-time event sources.

## Highlights

- Use-case and Medallion architecture design
- Implementation of Spark jobs and Airflow DAGs
- Reproducible integration of Kafka, SQL Server, MinIO and Delta Lake

## Results

- Ingestion patterns: 3. Fuente: docs/ejecucion-y-pruebas.md
- Data layers: 3. Fuente: docs/arquitectura-pmd.md

## Technology stack

Python · Apache Spark · Apache Airflow · Apache Kafka · Delta Lake · MinIO · SQL Server · Docker Compose

## Verifiable links

- Portfolio page: https://alonsomarcosm.github.io/en/projects/big-data-catalog-batch-streaming/
- GitHub repository: https://github.com/AlonsoMarcosM/spark-kafka-airflow-data-platform
- technical_docs: https://alonsomarcosm.github.io/spark-kafka-airflow-data-platform/
- architecture: https://alonsomarcosm.github.io/spark-kafka-airflow-data-platform/docs/visualizaciones/arquitectura-ejecutiva.svg
