Customer docs
Data Engineering Workspace Preview
A preview of a future private Ubuntu VM workspace for data engineering practice and small demos.
What it is
The Data Engineering Workspace is a future GangaCloud beta idea: a private Ubuntu VM sized for learning PySpark, JupyterLab, and small CSV or Parquet ETL demos.
It is a feasibility preview, not a managed Databricks replacement and not production Spark.
- Private Ubuntu VM for data engineering practice
- SSH ProxyJump access over private networking
- JupyterLab access through an SSH tunnel
- Recommended starter resources: 4 vCPU / 8 GB RAM / 80 GB disk
What customers can do
- Practice PySpark local mode on one VM
- Run JupyterLab privately through an SSH tunnel
- Learn CSV and Parquet ETL workflows
- Test small pandas, pyarrow, and PySpark examples
- Build private data engineering demos before moving to larger systems
Current limits
This preview is for learning and small demos only. It does not provide multi-node Spark, Kubernetes, Airflow, autoscaling, public IPv4, a production SLA, or a multi-tenant notebook service.
- No public IPv4 by default
- SSH tunnel required for browser tools
- Not suitable for sensitive production data yet
- Customers must review generated commands, SQL, and code before use