Skip to content

Repository files navigation

🎵 Spotify Listening Analytics - Production ML Infrastructure

Modern data engineering and machine learning infrastructure using Apache Spark, Delta Lake, MinIO, Trino, and Superset.

Python Spark Delta Lake Docker


🎯 Project Overview

A production-ready data lake analyzing Spotify listening patterns with 5 types of analytics (Descriptive, Diagnostic, Predictive, Prescriptive, Cognitive) powered by Apache Spark, Delta Lake, and MinIO object storage.

Key Features:

  • 🏗️ Medallion Architecture (Bronze → Silver → Gold)
  • 🗄️ S3-compatible object storage (MinIO)
  • 🤖 Machine Learning models (Valence & Energy prediction)
  • 📊 28 queryable analytics tables
  • 🎨 Interactive Superset dashboards
  • ⚡ Automated pipeline orchestration

⚡ Quick Start (3 Minutes)

# 1. Start all services
docker-compose up -d

# 2. Wait for services to be healthy (60 seconds)
docker-compose ps

# 3. Run full pipeline
docker-compose run --rm spotify-pipeline python3 run_full_pipeline.py

# 4. Query data with Trino
docker exec -it trino trino
USE delta_minio.default;
SELECT * FROM model_comparison;

# 5. Access dashboards
# Superset: http://localhost:8088 (admin/admin)
# MinIO Console: http://localhost:9001 (minioadmin/minioadmin123)

🏗️ Architecture

System Architecture Diagram

SPOTIFY API          KAGGLE DATASET (114K tracks)
     │                        │
     ▼                        ▼
┌─────────────────────────────────────────────┐
│   BRONZE LAYER (Raw Data)                   │
│   ├── listening_history_bronze              │
│   ├── my_tracks_features_bronze             │
│   ├── kaggle_tracks_bronze                  │
│   └── my_tracks_features_bronze_synthetic   │
└──────────────┬──────────────────────────────┘
               │ Apache Spark
               ▼
┌─────────────────────────────────────────────┐
│   SILVER LAYER (Cleaned & Enriched)         │
│   └── listening_with_features (5,926 rows)  │
│       ├── Audio features enrichment         │
│       ├── Temporal dimensions               │
│       └── Data quality validation           │
└──────────────┬──────────────────────────────┘
               │ Apache Spark + MLlib
               ▼
┌─────────────────────────────────────────────┐
│   GOLD LAYER (23 Analytics Tables)          │
│   ┌──────────────────────────────────────┐  │
│   │ Descriptive   │ Diagnostic           │  │
│   │ Predictive    │ Prescriptive         │  │
│   │ Cognitive     │                      │  │
│   └──────────────────────────────────────┘  │
└──────────────┬──────────────────────────────┘
               │
    ┌──────────┴───────────┐
    ▼                      ▼
┌─────────┐          ┌──────────┐
│  MinIO  │◄────────►│  Trino   │
│ Storage │  S3A     │   SQL    │
└─────────┘          └─────┬────┘
                           │
                           ▼
                     ┌──────────┐
                     │ Superset │
                     │Dashboards│
                     └──────────┘

Technology Stack

Component Technology Purpose
Data Ingestion Spotify Web API Real listening history
Processing Engine Apache Spark 3.5.3 Distributed data processing
Storage Format Delta Lake 3.2.1 ACID transactions on object storage
Object Storage MinIO (S3-compatible) Scalable, cloud-native storage
Query Engine Trino 435 Distributed SQL analytics
ML Framework Spark MLlib + scikit-learn Model training & validation
Visualization Apache Superset 3.1.0 Interactive dashboards
Orchestration Docker Compose + Scheduler Automated pipeline execution

📊 Data Layers & Tables

Storage Architecture

Dual-write mode (default): Data written to both local filesystem and MinIO S3-compatible storage

  • Bronze Bucket (spotify-bronze): 343MB raw data
  • Silver Bucket (spotify-silver): 9.3MB enriched data
  • Gold Bucket (spotify-gold): 3.3MB analytics tables
  • Total: 355MB across 4,716 Delta Lake objects

Bronze Layer (4 Tables)

Table Rows Description
listening_history_bronze 5,463 Personal Spotify listening history
my_tracks_features_bronze Variable Audio features from Spotify API
kaggle_tracks_bronze 114,000 Kaggle Spotify tracks dataset
my_tracks_features_bronze_synthetic Variable Synthetic features (API fallback)

Silver Layer (1 Table)

Table Rows Description
listening_with_features 5,926 Enriched listening data with audio features

Gold Layer (23 Tables)

Descriptive Analytics (5 tables) - What happened?

  • listening_patterns_by_time - Hourly/daily listening patterns
  • top_tracks_by_mood - Most-played tracks by mood
  • audio_feature_distributions - Feature distribution analysis
  • temporal_trends - Weekly and seasonal patterns
  • feature_source_coverage - Data source quality metrics

Diagnostic Analytics (5 tables) - Why did it happen?

  • mood_time_correlations - Time-of-day mood correlations
  • feature_correlations - Audio feature relationship matrix
  • weekend_vs_weekday - Behavioral pattern differences
  • mood_shift_patterns - Mood transition analysis
  • part_of_day_drivers - Key drivers by time of day

Predictive Analytics (5 tables) - What will happen?

  • mood_predictions - Valence forecasting (RMSE: 0.1947)
  • energy_forecasts - Energy prediction (RMSE: 0.2202)
  • model_comparison - Algorithm performance (6 models)
  • valence_model_comparison - Valence-specific model metrics
  • energy_model_comparison - Energy-specific model metrics

Prescriptive Analytics (4 tables) - What should we do?

  • mood_improvement_recommendations - Personalized playlists
  • optimal_listening_times - Best times for different moods
  • personalized_playlist_suggestions - Track recommendations
  • mood_intervention_triggers - When to intervene

Cognitive Analytics (4 tables) - Complex pattern discovery

  • mood_state_clusters - K-Means clustering (k=5, Silhouette: 0.2516)
  • listening_anomalies - Anomaly detection via Z-scores
  • sequential_patterns - Listening sequence analysis
  • behavioral_segments - User behavior archetypes

🤖 Machine Learning Models

Model Comparison Framework

3 Algorithms Trained & Compared:

Model Valence R² Energy R² Best Hyperparameters
Random Forest 0.134 0.197 numTrees=20, maxDepth=5, minInstancesPerNode=1
Gradient Boosting 0.182 0.205 maxIter=20, stepSize=0.1
Linear Regression 0.121 0.089 maxIter=100

Hyperparameter Tuning

# GridSearchCV with 3-fold cross-validation
param_grid = ParamGridBuilder() \
    .addGrid(rf.numTrees, [20, 50, 100]) \
    .addGrid(rf.maxDepth, [5, 10, 15]) \
    .addGrid(rf.minInstancesPerNode, [1, 5, 10]) \
    .build()  # 27 combinations tested

cv = CrossValidator(
    estimator=pipeline,
    estimatorParamMaps=param_grid,
    evaluator=RegressionEvaluator(metricName='rmse'),
    numFolds=3
)

best_model = cv.fit(train).bestModel

Cluster Validation (K-Means)

Rigorous methodology:

  • ✅ Elbow method (k=2 to k=8)
  • ✅ Silhouette analysis (optimal k=5, score=0.2516)
  • ✅ Stability testing (3 runs, std < 0.10)

Discovered 5 distinct mood states:

  • Cluster 0: High energy, moderate valence (142 songs)
  • Cluster 1: Balanced mood - largest group (390 songs)
  • Cluster 2: Low valence/energy - melancholic (318 songs)
  • Cluster 3: Happy & energetic (368 songs)
  • Cluster 4: Mid-tempo danceable (286 songs)

🗄️ MinIO Object Storage

Why MinIO?

  • ✅ Scalable: Handle petabytes of data (2.6Tbps reads, 1.32Tbps writes)
  • ✅ Cloud-Native: S3-compatible, works with AWS/GCP/Azure
  • ✅ Production-Ready: ACID transactions via Delta Lake
  • ✅ Cost-Effective: Free open-source alternative
  • ✅ Multi-User: Shared storage for concurrent access

Storage Modes

Configure via STORAGE_MODE environment variable:

Mode Description Use Case
local Local filesystem only Development, single machine
minio MinIO object storage only Production, cloud deployments
dual Write to both (default) Safe migration, validation

MinIO Buckets

# Access MinIO Console: http://localhost:9001
# Login: minioadmin / minioadmin123

Buckets:
├── spotify-bronze/     # 343MB raw data
├── spotify-silver/     # 9.3MB enriched data
├── spotify-gold/       # 3.3MB analytics
└── spotify-exports/    # CSV exports

Query MinIO Data via Trino

-- Connect to Trino
docker exec -it trino trino

-- Use MinIO catalog
USE delta_minio.default;

-- Show all tables
SHOW TABLES;

-- Query listening patterns
SELECT hour_of_day, day_name, play_count, avg_valence, avg_energy
FROM listening_patterns_by_time
ORDER BY play_count DESC
LIMIT 10;

-- Check ML model performance
SELECT model_name, test_r2, test_rmse, training_time_sec
FROM model_comparison
ORDER BY test_r2 DESC;

-- Explore mood clusters
SELECT cluster, member_count, avg_valence, avg_energy
FROM mood_state_clusters;

🚀 Running the Pipeline

Full Pipeline (All Stages)

# Runs: Bronze → Silver → Gold → ML (5-7 minutes)
docker-compose run --rm spotify-pipeline python3 run_full_pipeline.py

Pipeline Stages:

  1. Bronze Ingestion: Spotify API + Kaggle dataset (2 min)
  2. Silver Enrichment: Feature engineering + validation (1 min)
  3. Gold Descriptive: Listening patterns analysis (30 sec)
  4. Gold Diagnostic: Correlation analysis (30 sec)
  5. Gold Predictive: ML model training (2 min)
  6. Gold Prescriptive: Recommendation generation (30 sec)
  7. Gold Cognitive: Clustering + anomaly detection (1 min)

Individual Stages

# Bronze layer only
docker-compose run --rm spotify-pipeline python3 run_ingestion.py

# Silver layer only
docker-compose run --rm spotify-pipeline \
  python3 scripts/build_silver_listening_with_features.py

# Gold descriptive analytics
docker-compose run --rm spotify-pipeline \
  python3 gold/descriptive/build_descriptive_analytics.py

# Predictive ML models
docker-compose run --rm spotify-pipeline \
  python3 gold/predictive/build_predictive_models.py

# Cognitive analytics (clustering)
docker-compose run --rm spotify-pipeline \
  python3 gold/cognitive/build_cognitive_analytics.py

Automated Scheduling

The pipeline runs automatically every 6 hours via scheduler.py:

# Check scheduler logs
docker-compose logs spotify-scheduler --tail=100

🎨 Superset Dashboards

Access Superset

# URL: http://localhost:8088
# Username: admin
# Password: admin

Creating Dashboards

See SUPERSET_DASHBOARD_GUIDE.md for detailed instructions on creating:

  1. ML Model Performance Dashboard

    • Model comparison table
    • Performance metrics (R², RMSE)
    • Hyperparameter visualization
    • Training time analysis
  2. Listening Patterns Dashboard

    • Hourly/daily listening heatmap
    • Weekend vs weekday comparison
    • Top tracks by mood
    • Temporal trends
  3. Cluster Analysis Dashboard

    • Elbow curve for optimal k
    • Silhouette score visualization
    • Cluster profile table
    • Mood state distribution
  4. Prescriptive Analytics Dashboard

    • Personalized playlist recommendations
    • Optimal listening times
    • Mood improvement suggestions

📋 System Requirements

Prerequisites

  • Docker Desktop (with Docker Compose v2+)
  • RAM: 8GB minimum, 16GB recommended
  • Disk Space: 20GB free space
  • OS: Windows 10+, macOS 10.15+, or Linux

Services & Ports

Service Container Port Purpose
MinIO Server minio 9000 S3 API endpoint
MinIO Console minio 9001 Web UI
Trino trino 8080 Query engine
Superset superset 8088 Dashboards
PostgreSQL postgres 5432 Superset metadata
Redis redis 6379 Superset cache

🛠️ Installation

1. Clone Repository

git clone https://github.com/yourusername/spotify-analytics.git
cd spotify-analytics

2. Configure Spotify API (Optional)

# Create .env file
cp .env.example .env

# Edit with your credentials
# SPOTIFY_CLIENT_ID=your_client_id
# SPOTIFY_CLIENT_SECRET=your_client_secret
# SPOTIFY_REDIRECT_URI=http://localhost:8888/callback

3. Start Services

# Start all containers
docker-compose up -d

# Verify all services are healthy
docker-compose ps

# Expected output:
# minio        Up 2 minutes (healthy)
# trino        Up 2 minutes (healthy)
# superset     Up 2 minutes (healthy)
# postgres     Up 2 minutes (healthy)
# redis        Up 2 minutes (healthy)
# scheduler    Up 2 minutes

4. Register Delta Tables in Trino

# Register all 28 tables with Trino
docker exec -i trino trino --execute "$(cat register_minio_tables.sql)"

# Verify tables registered
docker exec -i trino trino --execute "USE delta_minio.default; SHOW TABLES;"

5. Run Pipeline

# First run (populates all layers)
docker-compose run --rm spotify-pipeline python3 run_full_pipeline.py

# Check execution logs
docker-compose logs spotify-scheduler

📖 Project Structure

spotify-analytics/
├── bronze/                      # Bronze layer ingestion scripts
│   ├── ingest_listening_history.py
│   ├── ingest_kaggle_tracks.py
│   └── ingest_my_tracks_features.py
├── config/                      # Configuration
│   ├── settings.py
│   └── minio_config.py         # Storage configuration
├── gold/                        # Gold layer analytics
│   ├── descriptive/
│   ├── diagnostic/
│   ├── predictive/
│   ├── prescriptive/
│   └── cognitive/
├── scripts/                     # Helper scripts
│   ├── build_silver_listening_with_features.py
│   ├── export_gold_to_csv.py
│   └── populate_missing_features.py
├── trino/                       # Trino configuration
│   └── catalog/
│       ├── delta.properties
│       └── delta_minio.properties
├── utils/                       # Utilities
│   └── spark_helper.py         # Spark session factory
├── writers/                     # Data writers
│   ├── delta_writer.py
│   └── unified_delta_writer.py
├── docker-compose.yml           # Service orchestration
├── Dockerfile                   # Spark container image
├── run_full_pipeline.py         # Main pipeline script
├── run_ingestion.py            # Bronze layer ingestion
├── scheduler.py                # Automated scheduling
├── register_minio_tables.sql   # Trino table registration
└── README.md                   # This file

🔍 Example Queries

Listening Patterns

-- Peak listening times
SELECT hour_of_day, day_name, play_count, avg_valence, avg_energy
FROM listening_patterns_by_time
WHERE play_count > 200
ORDER BY play_count DESC;

-- Weekend vs weekday differences
SELECT
  is_weekend,
  AVG(play_count) as avg_plays,
  AVG(avg_valence) as avg_valence,
  AVG(avg_energy) as avg_energy
FROM listening_patterns_by_time
GROUP BY is_weekend;

Machine Learning Results

-- Model performance comparison
SELECT model_name, task, test_r2, test_rmse, cv_mean_rmse, training_time_sec
FROM model_comparison
ORDER BY test_r2 DESC;

-- Hyperparameter configurations
SELECT model_name, best_params
FROM model_comparison
WHERE test_r2 > 0.15;

Prescriptive Recommendations

-- Personalized playlist suggestions
SELECT playlist_name, playlist_purpose, track_name, artist_name
FROM personalized_playlist_suggestions
WHERE playlist_name = 'Focus & Deep Work'
LIMIT 10;

-- Optimal listening times by mood
SELECT time_window, recommended_valence_range, avg_historical_valence
FROM optimal_listening_times
ORDER BY time_window;

Cluster Analysis

-- Mood cluster profiles
SELECT cluster, member_count, avg_valence, avg_energy,
       avg_danceability, avg_acousticness
FROM mood_state_clusters
ORDER BY member_count DESC;

📚 Documentation

Document Description
README.md Main project documentation (this file)
SUPERSET_DASHBOARD_GUIDE.md Step-by-step dashboard creation guide
MINIO_INTEGRATION_GUIDE.md MinIO setup, configuration, and troubleshooting

🎯 Key Features

Data Engineering

✅ Medallion Architecture

  • Bronze: Raw data preservation with append-only semantics
  • Silver: Quality-assured, enriched data with deduplication
  • Gold: Business-ready analytics tables

✅ Delta Lake ACID Guarantees

  • Atomic operations
  • Time-travel queries
  • Schema evolution
  • Concurrent read/write support

✅ Data Quality Framework

  • Automated validation rules
  • Null value handling
  • Feature range constraints
  • Outlier detection

✅ Dual-Storage Strategy

  • Local filesystem + MinIO object storage
  • Zero-downtime migration capability
  • Data consistency validation

Machine Learning

✅ Hyperparameter Optimization

  • GridSearchCV with 27+ parameter combinations
  • 3-fold cross-validation
  • Parallel execution
  • Best model selection

✅ Model Comparison

  • Multiple algorithms (RF, GBT, LinearReg)
  • Unified evaluation framework
  • Baseline comparisons
  • Results stored in Gold layer

✅ Rigorous Validation

  • Train/test split (80/20)
  • Cross-validation for honest estimates
  • Multiple metrics (R², RMSE, MAE)
  • Cluster stability testing

DevOps & Automation

✅ Docker Orchestration

  • 6 containerized services
  • Automated dependency management
  • Health checks for all services
  • One-command deployment

✅ Automated Pipelines

  • Scheduled execution (every 6 hours)
  • Bronze → Silver → Gold → ML workflow
  • Comprehensive logging
  • Error handling and retries

📊 Project Metrics

Metric Value
Total Tables 28 (4 Bronze + 1 Silver + 23 Gold)
Total Data Volume 355MB (MinIO storage)
Total Objects 4,716 Delta Lake objects
ML Models 6 trained models (3 algorithms × 2 targets)
Cluster Analysis 5 mood clusters identified
Pipeline Duration ~5-7 minutes (full run)
Automation 100% automated (Bronze → Gold)
Services 6 Docker containers
Code Lines ~3,500 Python/SQL

⚠️ Current Limitations

Synthetic Audio Features

Status: Using deterministic synthetic features due to Spotify API restrictions

Impact:

  • ✅ Listening history is 100% real (timestamps, tracks, artists)
  • ✅ Infrastructure is production-ready
  • ✅ ML methodology is valid
  • ⚠️ ML predictions demonstrate methodology only (R² < 0.25 on synthetic data)

Future: Request extended Spotify API quota for real audio features


🔧 Troubleshooting

MinIO Connection Issues

# Check MinIO status
docker-compose ps minio

# View MinIO logs
docker-compose logs minio --tail=50

# Restart MinIO
docker-compose restart minio

Trino Not Querying Tables

# Re-register tables
docker exec -i trino trino --execute "$(cat register_minio_tables.sql)"

# Check Trino catalog
docker exec -it trino trino
SHOW CATALOGS;
USE delta_minio.default;
SHOW TABLES;

Pipeline Failures

# Check scheduler logs
docker-compose logs spotify-scheduler --tail=100

# Run pipeline manually with verbose output
docker-compose run --rm spotify-pipeline python3 run_full_pipeline.py

Out of Memory Errors

# Increase Docker memory allocation
# Docker Desktop → Settings → Resources → Memory: 8GB+

# Restart containers
docker-compose restart

🤝 Contributing

Contributions welcome! Areas for enhancement:

  • Real-time streaming with Apache Kafka
  • Additional ML algorithms (XGBoost, LightGBM)
  • Model explainability (SHAP values)
  • Advanced anomaly detection
  • A/B testing framework
  • Real-time dashboard updates

📄 License

This project is for educational and portfolio demonstration purposes.

Data Sources:


🙏 Acknowledgments

Technologies:

Datasets:


📧 Contact & Links

LinkedIn: [in/sambett] Email: [sbettaie56@gmail.com]


Last Updated: 2025-12-04 Status: ✅ Production-Ready | 🗄️ MinIO Integrated | 🤖 28 Tables Queryable | 📊 All 5 Analytics Types Complete

About

Batch data lake for Spotify listening analytics — Bronze/Silver/Gold medallion pipeline on Apache Spark + Delta Lake, MinIO object storage, Trino SQL, Superset dashboards, orchestrated via Docker Compose.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages