Final project for a university course on 'Big Data.' The project's objective is to compare various GDBMS solutions in terms of usability and time performance within the domain of GitHub data analysis. This involves defining different analytical scenarios and utilizing different batches of the original Big Query dataset. The entire project is built within a Dockerized environment to ensure isolation and the potential for future scalability to a cluster of nodes with fault tolerance. Additionally, this project incorporates several technologies, including Spark for preprocessing, Neo4j, TigerGraph, and ArangoDB.
This repository is directly forked and inspired from Big Data Europe repositories
Docker Compose containing:
- Apache Spark cluster running one Spark Master and multiple Spark workers
- Hadoop HDFS cluster
- Neo4j ™ Neo4j graph database soluction 1
- TigerGraph ™ TigerGraph graph database soluction 2
- ArangoDB ™ ArangoDB graph database soluction 2
- Jupyter Lab service to test PySpark jobs
To start the docker big data playground repository:
cd docker-bigdata-playground
docker-compose up
Move the dataset files into the path /datasets.
Log into the container and put the file into HDFS:
$ docker-compose exec spark-master bash
> ./scripts/exec-load-data.sh
Connect to Jupyter Hub by accessing container logs:
$ docker logs jupyter-notebooks
> 2023-05-03 17:29:39 To access the server, open this file in a browser:
> 2023-05-03 17:29:39 file:///home/jovyan/.local/share/jupyter/runtime/jpserver-7-open.html
> 2023-05-03 17:29:39 Or copy and paste one of these URLs:
> 2023-05-03 17:29:39 http://083d9da0d714:8888/lab?token=686167f3cee298e578315d50990c397ffd09b75cb5705cf3
> 2023-05-03 17:29:39 or http://127.0.0.1:8888/lab?token=686167f3cee298e578315d50990c397ffd09b75cb5705cf3
Click on the last line in the logs, enter Jupyter Hub in your brower and follow instructions in the notebooks/BigDataPipeline.ipynb













