GangaCloud

Customer docs

Data Engineering Workspace Preview

A preview of a future private Ubuntu VM workspace for data engineering practice and small demos.

What it is

The Data Engineering Workspace is a future GangaCloud beta idea: a private Ubuntu VM sized for learning PySpark, JupyterLab, and small CSV or Parquet ETL demos.

It is a feasibility preview, not a managed Databricks replacement and not production Spark.

  • Private Ubuntu VM for data engineering practice
  • SSH ProxyJump access over private networking
  • JupyterLab access through an SSH tunnel
  • Recommended starter resources: 4 vCPU / 8 GB RAM / 80 GB disk

What customers can do

  • Practice PySpark local mode on one VM
  • Run JupyterLab privately through an SSH tunnel
  • Learn CSV and Parquet ETL workflows
  • Test small pandas, pyarrow, and PySpark examples
  • Build private data engineering demos before moving to larger systems

Current limits

This preview is for learning and small demos only. It does not provide multi-node Spark, Kubernetes, Airflow, autoscaling, public IPv4, a production SLA, or a multi-tenant notebook service.

  • No public IPv4 by default
  • SSH tunnel required for browser tools
  • Not suitable for sensitive production data yet
  • Customers must review generated commands, SQL, and code before use