Modern data engineering and machine learning infrastructure using Apache Spark, Delta Lake, MinIO, Trino, and Superset.
A production-ready data lake analyzing Spotify listening patterns with 5 types of analytics (Descriptive, Diagnostic, Predictive, Prescriptive, Cognitive) powered by Apache Spark, Delta Lake, and MinIO object storage.
Key Features:
- 🏗️ Medallion Architecture (Bronze → Silver → Gold)
- 🗄️ S3-compatible object storage (MinIO)
- 🤖 Machine Learning models (Valence & Energy prediction)
- 📊 28 queryable analytics tables
- 🎨 Interactive Superset dashboards
- ⚡ Automated pipeline orchestration
# 1. Start all services
docker-compose up -d
# 2. Wait for services to be healthy (60 seconds)
docker-compose ps
# 3. Run full pipeline
docker-compose run --rm spotify-pipeline python3 run_full_pipeline.py
# 4. Query data with Trino
docker exec -it trino trino
USE delta_minio.default;
SELECT * FROM model_comparison;
# 5. Access dashboards
# Superset: http://localhost:8088 (admin/admin)
# MinIO Console: http://localhost:9001 (minioadmin/minioadmin123)SPOTIFY API KAGGLE DATASET (114K tracks)
│ │
▼ ▼
┌─────────────────────────────────────────────┐
│ BRONZE LAYER (Raw Data) │
│ ├── listening_history_bronze │
│ ├── my_tracks_features_bronze │
│ ├── kaggle_tracks_bronze │
│ └── my_tracks_features_bronze_synthetic │
└──────────────┬──────────────────────────────┘
│ Apache Spark
▼
┌─────────────────────────────────────────────┐
│ SILVER LAYER (Cleaned & Enriched) │
│ └── listening_with_features (5,926 rows) │
│ ├── Audio features enrichment │
│ ├── Temporal dimensions │
│ └── Data quality validation │
└──────────────┬──────────────────────────────┘
│ Apache Spark + MLlib
▼
┌─────────────────────────────────────────────┐
│ GOLD LAYER (23 Analytics Tables) │
│ ┌──────────────────────────────────────┐ │
│ │ Descriptive │ Diagnostic │ │
│ │ Predictive │ Prescriptive │ │
│ │ Cognitive │ │ │
│ └──────────────────────────────────────┘ │
└──────────────┬──────────────────────────────┘
│
┌──────────┴───────────┐
▼ ▼
┌─────────┐ ┌──────────┐
│ MinIO │◄────────►│ Trino │
│ Storage │ S3A │ SQL │
└─────────┘ └─────┬────┘
│
▼
┌──────────┐
│ Superset │
│Dashboards│
└──────────┘
| Component | Technology | Purpose |
|---|---|---|
| Data Ingestion | Spotify Web API | Real listening history |
| Processing Engine | Apache Spark 3.5.3 | Distributed data processing |
| Storage Format | Delta Lake 3.2.1 | ACID transactions on object storage |
| Object Storage | MinIO (S3-compatible) | Scalable, cloud-native storage |
| Query Engine | Trino 435 | Distributed SQL analytics |
| ML Framework | Spark MLlib + scikit-learn | Model training & validation |
| Visualization | Apache Superset 3.1.0 | Interactive dashboards |
| Orchestration | Docker Compose + Scheduler | Automated pipeline execution |
Dual-write mode (default): Data written to both local filesystem and MinIO S3-compatible storage
- Bronze Bucket (
spotify-bronze): 343MB raw data - Silver Bucket (
spotify-silver): 9.3MB enriched data - Gold Bucket (
spotify-gold): 3.3MB analytics tables - Total: 355MB across 4,716 Delta Lake objects
| Table | Rows | Description |
|---|---|---|
listening_history_bronze |
5,463 | Personal Spotify listening history |
my_tracks_features_bronze |
Variable | Audio features from Spotify API |
kaggle_tracks_bronze |
114,000 | Kaggle Spotify tracks dataset |
my_tracks_features_bronze_synthetic |
Variable | Synthetic features (API fallback) |
| Table | Rows | Description |
|---|---|---|
listening_with_features |
5,926 | Enriched listening data with audio features |
Descriptive Analytics (5 tables) - What happened?
listening_patterns_by_time- Hourly/daily listening patternstop_tracks_by_mood- Most-played tracks by moodaudio_feature_distributions- Feature distribution analysistemporal_trends- Weekly and seasonal patternsfeature_source_coverage- Data source quality metrics
Diagnostic Analytics (5 tables) - Why did it happen?
mood_time_correlations- Time-of-day mood correlationsfeature_correlations- Audio feature relationship matrixweekend_vs_weekday- Behavioral pattern differencesmood_shift_patterns- Mood transition analysispart_of_day_drivers- Key drivers by time of day
Predictive Analytics (5 tables) - What will happen?
mood_predictions- Valence forecasting (RMSE: 0.1947)energy_forecasts- Energy prediction (RMSE: 0.2202)model_comparison- Algorithm performance (6 models)valence_model_comparison- Valence-specific model metricsenergy_model_comparison- Energy-specific model metrics
Prescriptive Analytics (4 tables) - What should we do?
mood_improvement_recommendations- Personalized playlistsoptimal_listening_times- Best times for different moodspersonalized_playlist_suggestions- Track recommendationsmood_intervention_triggers- When to intervene
Cognitive Analytics (4 tables) - Complex pattern discovery
mood_state_clusters- K-Means clustering (k=5, Silhouette: 0.2516)listening_anomalies- Anomaly detection via Z-scoressequential_patterns- Listening sequence analysisbehavioral_segments- User behavior archetypes
3 Algorithms Trained & Compared:
| Model | Valence R² | Energy R² | Best Hyperparameters |
|---|---|---|---|
| Random Forest | 0.134 | 0.197 | numTrees=20, maxDepth=5, minInstancesPerNode=1 |
| Gradient Boosting | 0.182 | 0.205 | maxIter=20, stepSize=0.1 |
| Linear Regression | 0.121 | 0.089 | maxIter=100 |
# GridSearchCV with 3-fold cross-validation
param_grid = ParamGridBuilder() \
.addGrid(rf.numTrees, [20, 50, 100]) \
.addGrid(rf.maxDepth, [5, 10, 15]) \
.addGrid(rf.minInstancesPerNode, [1, 5, 10]) \
.build() # 27 combinations tested
cv = CrossValidator(
estimator=pipeline,
estimatorParamMaps=param_grid,
evaluator=RegressionEvaluator(metricName='rmse'),
numFolds=3
)
best_model = cv.fit(train).bestModelRigorous methodology:
- ✅ Elbow method (k=2 to k=8)
- ✅ Silhouette analysis (optimal k=5, score=0.2516)
- ✅ Stability testing (3 runs, std < 0.10)
Discovered 5 distinct mood states:
- Cluster 0: High energy, moderate valence (142 songs)
- Cluster 1: Balanced mood - largest group (390 songs)
- Cluster 2: Low valence/energy - melancholic (318 songs)
- Cluster 3: Happy & energetic (368 songs)
- Cluster 4: Mid-tempo danceable (286 songs)
- ✅ Scalable: Handle petabytes of data (2.6Tbps reads, 1.32Tbps writes)
- ✅ Cloud-Native: S3-compatible, works with AWS/GCP/Azure
- ✅ Production-Ready: ACID transactions via Delta Lake
- ✅ Cost-Effective: Free open-source alternative
- ✅ Multi-User: Shared storage for concurrent access
Configure via STORAGE_MODE environment variable:
| Mode | Description | Use Case |
|---|---|---|
local |
Local filesystem only | Development, single machine |
minio |
MinIO object storage only | Production, cloud deployments |
dual |
Write to both (default) | Safe migration, validation |
# Access MinIO Console: http://localhost:9001
# Login: minioadmin / minioadmin123
Buckets:
├── spotify-bronze/ # 343MB raw data
├── spotify-silver/ # 9.3MB enriched data
├── spotify-gold/ # 3.3MB analytics
└── spotify-exports/ # CSV exports-- Connect to Trino
docker exec -it trino trino
-- Use MinIO catalog
USE delta_minio.default;
-- Show all tables
SHOW TABLES;
-- Query listening patterns
SELECT hour_of_day, day_name, play_count, avg_valence, avg_energy
FROM listening_patterns_by_time
ORDER BY play_count DESC
LIMIT 10;
-- Check ML model performance
SELECT model_name, test_r2, test_rmse, training_time_sec
FROM model_comparison
ORDER BY test_r2 DESC;
-- Explore mood clusters
SELECT cluster, member_count, avg_valence, avg_energy
FROM mood_state_clusters;# Runs: Bronze → Silver → Gold → ML (5-7 minutes)
docker-compose run --rm spotify-pipeline python3 run_full_pipeline.pyPipeline Stages:
- Bronze Ingestion: Spotify API + Kaggle dataset (2 min)
- Silver Enrichment: Feature engineering + validation (1 min)
- Gold Descriptive: Listening patterns analysis (30 sec)
- Gold Diagnostic: Correlation analysis (30 sec)
- Gold Predictive: ML model training (2 min)
- Gold Prescriptive: Recommendation generation (30 sec)
- Gold Cognitive: Clustering + anomaly detection (1 min)
# Bronze layer only
docker-compose run --rm spotify-pipeline python3 run_ingestion.py
# Silver layer only
docker-compose run --rm spotify-pipeline \
python3 scripts/build_silver_listening_with_features.py
# Gold descriptive analytics
docker-compose run --rm spotify-pipeline \
python3 gold/descriptive/build_descriptive_analytics.py
# Predictive ML models
docker-compose run --rm spotify-pipeline \
python3 gold/predictive/build_predictive_models.py
# Cognitive analytics (clustering)
docker-compose run --rm spotify-pipeline \
python3 gold/cognitive/build_cognitive_analytics.pyThe pipeline runs automatically every 6 hours via scheduler.py:
# Check scheduler logs
docker-compose logs spotify-scheduler --tail=100# URL: http://localhost:8088
# Username: admin
# Password: adminSee SUPERSET_DASHBOARD_GUIDE.md for detailed instructions on creating:
-
ML Model Performance Dashboard
- Model comparison table
- Performance metrics (R², RMSE)
- Hyperparameter visualization
- Training time analysis
-
Listening Patterns Dashboard
- Hourly/daily listening heatmap
- Weekend vs weekday comparison
- Top tracks by mood
- Temporal trends
-
Cluster Analysis Dashboard
- Elbow curve for optimal k
- Silhouette score visualization
- Cluster profile table
- Mood state distribution
-
Prescriptive Analytics Dashboard
- Personalized playlist recommendations
- Optimal listening times
- Mood improvement suggestions
- Docker Desktop (with Docker Compose v2+)
- RAM: 8GB minimum, 16GB recommended
- Disk Space: 20GB free space
- OS: Windows 10+, macOS 10.15+, or Linux
| Service | Container | Port | Purpose |
|---|---|---|---|
| MinIO Server | minio |
9000 | S3 API endpoint |
| MinIO Console | minio |
9001 | Web UI |
| Trino | trino |
8080 | Query engine |
| Superset | superset |
8088 | Dashboards |
| PostgreSQL | postgres |
5432 | Superset metadata |
| Redis | redis |
6379 | Superset cache |
git clone https://github.com/yourusername/spotify-analytics.git
cd spotify-analytics# Create .env file
cp .env.example .env
# Edit with your credentials
# SPOTIFY_CLIENT_ID=your_client_id
# SPOTIFY_CLIENT_SECRET=your_client_secret
# SPOTIFY_REDIRECT_URI=http://localhost:8888/callback# Start all containers
docker-compose up -d
# Verify all services are healthy
docker-compose ps
# Expected output:
# minio Up 2 minutes (healthy)
# trino Up 2 minutes (healthy)
# superset Up 2 minutes (healthy)
# postgres Up 2 minutes (healthy)
# redis Up 2 minutes (healthy)
# scheduler Up 2 minutes# Register all 28 tables with Trino
docker exec -i trino trino --execute "$(cat register_minio_tables.sql)"
# Verify tables registered
docker exec -i trino trino --execute "USE delta_minio.default; SHOW TABLES;"# First run (populates all layers)
docker-compose run --rm spotify-pipeline python3 run_full_pipeline.py
# Check execution logs
docker-compose logs spotify-schedulerspotify-analytics/
├── bronze/ # Bronze layer ingestion scripts
│ ├── ingest_listening_history.py
│ ├── ingest_kaggle_tracks.py
│ └── ingest_my_tracks_features.py
├── config/ # Configuration
│ ├── settings.py
│ └── minio_config.py # Storage configuration
├── gold/ # Gold layer analytics
│ ├── descriptive/
│ ├── diagnostic/
│ ├── predictive/
│ ├── prescriptive/
│ └── cognitive/
├── scripts/ # Helper scripts
│ ├── build_silver_listening_with_features.py
│ ├── export_gold_to_csv.py
│ └── populate_missing_features.py
├── trino/ # Trino configuration
│ └── catalog/
│ ├── delta.properties
│ └── delta_minio.properties
├── utils/ # Utilities
│ └── spark_helper.py # Spark session factory
├── writers/ # Data writers
│ ├── delta_writer.py
│ └── unified_delta_writer.py
├── docker-compose.yml # Service orchestration
├── Dockerfile # Spark container image
├── run_full_pipeline.py # Main pipeline script
├── run_ingestion.py # Bronze layer ingestion
├── scheduler.py # Automated scheduling
├── register_minio_tables.sql # Trino table registration
└── README.md # This file
-- Peak listening times
SELECT hour_of_day, day_name, play_count, avg_valence, avg_energy
FROM listening_patterns_by_time
WHERE play_count > 200
ORDER BY play_count DESC;
-- Weekend vs weekday differences
SELECT
is_weekend,
AVG(play_count) as avg_plays,
AVG(avg_valence) as avg_valence,
AVG(avg_energy) as avg_energy
FROM listening_patterns_by_time
GROUP BY is_weekend;-- Model performance comparison
SELECT model_name, task, test_r2, test_rmse, cv_mean_rmse, training_time_sec
FROM model_comparison
ORDER BY test_r2 DESC;
-- Hyperparameter configurations
SELECT model_name, best_params
FROM model_comparison
WHERE test_r2 > 0.15;-- Personalized playlist suggestions
SELECT playlist_name, playlist_purpose, track_name, artist_name
FROM personalized_playlist_suggestions
WHERE playlist_name = 'Focus & Deep Work'
LIMIT 10;
-- Optimal listening times by mood
SELECT time_window, recommended_valence_range, avg_historical_valence
FROM optimal_listening_times
ORDER BY time_window;-- Mood cluster profiles
SELECT cluster, member_count, avg_valence, avg_energy,
avg_danceability, avg_acousticness
FROM mood_state_clusters
ORDER BY member_count DESC;| Document | Description |
|---|---|
| README.md | Main project documentation (this file) |
| SUPERSET_DASHBOARD_GUIDE.md | Step-by-step dashboard creation guide |
| MINIO_INTEGRATION_GUIDE.md | MinIO setup, configuration, and troubleshooting |
✅ Medallion Architecture
- Bronze: Raw data preservation with append-only semantics
- Silver: Quality-assured, enriched data with deduplication
- Gold: Business-ready analytics tables
✅ Delta Lake ACID Guarantees
- Atomic operations
- Time-travel queries
- Schema evolution
- Concurrent read/write support
✅ Data Quality Framework
- Automated validation rules
- Null value handling
- Feature range constraints
- Outlier detection
✅ Dual-Storage Strategy
- Local filesystem + MinIO object storage
- Zero-downtime migration capability
- Data consistency validation
✅ Hyperparameter Optimization
- GridSearchCV with 27+ parameter combinations
- 3-fold cross-validation
- Parallel execution
- Best model selection
✅ Model Comparison
- Multiple algorithms (RF, GBT, LinearReg)
- Unified evaluation framework
- Baseline comparisons
- Results stored in Gold layer
✅ Rigorous Validation
- Train/test split (80/20)
- Cross-validation for honest estimates
- Multiple metrics (R², RMSE, MAE)
- Cluster stability testing
✅ Docker Orchestration
- 6 containerized services
- Automated dependency management
- Health checks for all services
- One-command deployment
✅ Automated Pipelines
- Scheduled execution (every 6 hours)
- Bronze → Silver → Gold → ML workflow
- Comprehensive logging
- Error handling and retries
| Metric | Value |
|---|---|
| Total Tables | 28 (4 Bronze + 1 Silver + 23 Gold) |
| Total Data Volume | 355MB (MinIO storage) |
| Total Objects | 4,716 Delta Lake objects |
| ML Models | 6 trained models (3 algorithms × 2 targets) |
| Cluster Analysis | 5 mood clusters identified |
| Pipeline Duration | ~5-7 minutes (full run) |
| Automation | 100% automated (Bronze → Gold) |
| Services | 6 Docker containers |
| Code Lines | ~3,500 Python/SQL |
Status: Using deterministic synthetic features due to Spotify API restrictions
Impact:
- ✅ Listening history is 100% real (timestamps, tracks, artists)
- ✅ Infrastructure is production-ready
- ✅ ML methodology is valid
⚠️ ML predictions demonstrate methodology only (R² < 0.25 on synthetic data)
Future: Request extended Spotify API quota for real audio features
# Check MinIO status
docker-compose ps minio
# View MinIO logs
docker-compose logs minio --tail=50
# Restart MinIO
docker-compose restart minio# Re-register tables
docker exec -i trino trino --execute "$(cat register_minio_tables.sql)"
# Check Trino catalog
docker exec -it trino trino
SHOW CATALOGS;
USE delta_minio.default;
SHOW TABLES;# Check scheduler logs
docker-compose logs spotify-scheduler --tail=100
# Run pipeline manually with verbose output
docker-compose run --rm spotify-pipeline python3 run_full_pipeline.py# Increase Docker memory allocation
# Docker Desktop → Settings → Resources → Memory: 8GB+
# Restart containers
docker-compose restartContributions welcome! Areas for enhancement:
- Real-time streaming with Apache Kafka
- Additional ML algorithms (XGBoost, LightGBM)
- Model explainability (SHAP values)
- Advanced anomaly detection
- A/B testing framework
- Real-time dashboard updates
This project is for educational and portfolio demonstration purposes.
Data Sources:
- Personal Spotify data: Personal use only
- Kaggle dataset: Kaggle Terms of Use
Technologies:
- Apache Spark - Distributed processing
- Delta Lake - ACID storage layer
- MinIO - S3-compatible object storage
- Trino - Distributed SQL engine
- Apache Superset - Data visualization
Datasets:
- Spotify Tracks Dataset (Kaggle, 114K tracks)
LinkedIn: [in/sambett] Email: [sbettaie56@gmail.com]
Last Updated: 2025-12-04 Status: ✅ Production-Ready | 🗄️ MinIO Integrated | 🤖 28 Tables Queryable | 📊 All 5 Analytics Types Complete