User Guide & Demo

How to Run Dataproc 3.0 Locally with Docker

Published 6 min readData engineering

Spinning up a cloud Dataproc cluster just to test a PySpark transformation or validate an SQL query wastes minutes in provisioning and quickly racks up GCP charges. With LocalCloud Dataproc images, developers and AI coding agents can run the exact Dataproc 3.0 component stack entirely on their local machines — with zero cloud credentials, zero billing risk, and zero external dependencies.

For AI coding agents

Copy-paste prompt for Claude Code, Cursor, or Gemini CLI

Give this prompt to your AI coding agent to execute a safe local verification loop without touching real Google Cloud credentials:

Run a local Apache Spark smoke test using the agentcloud/dataproc:3.0.0-debian13 container.
Do not request or use GCP cloud credentials.
1. Run the baked-in self-test suite: docker run --rm agentcloud/dataproc:3.0.0-debian13 self-test
2. Execute a PySpark aggregation against local test data.
3. Report component versions and verification results.
Why it matters

The local feedback loop for developers and agents

For Developers

  • Zero spin-up lag: Skip the 90–180 second cloud cluster provisioning delay on every code change.
  • Zero cloud costs: Iterate, fail, debug, and profile Spark jobs locally without unexpected bills.
  • Offline portability: Develop pipelines on an airplane or local network without an active internet connection.

For AI Coding Agents

  • Safe sandbox boundary: Agents can generate, run, and benchmark PySpark code without needing IAM permissions.
  • Fast self-healing: If an agent writes invalid Spark SQL or bad PySpark code, it gets instant tracebacks to self-correct.
  • Deterministic builds: Digest-pinned public images prevent silent upstream drift across developer machines.
Component stack

What's inside Dataproc 3.0 (Debian 13)

The agentcloud/dataproc:3.0.0-debian13 container packages the complete software stack corresponding to Google Cloud Dataproc 3.0, built from verified public digests:

Component Version Installation Path Role
Apache Spark 4.1.2 /opt/localcloud/spark PySpark & Spark SQL execution
Apache Hadoop 3.5.0 /opt/localcloud/hadoop HDFS distributed storage & YARN scheduler
Apache Hive 4.2.0 /opt/localcloud/hive Metastore & HiveServer2 (pre-initialized Derby)
Java (OpenJDK) 21 (Temurin) /opt/localcloud/java JVM runtime with security entropy options
Python 3.12 (CPython) /opt/localcloud/python PySpark driver and executor environment
GCP Connectors GCS & BigQuery /opt/localcloud/spark/jars Native connectors for Cloud Storage & BigQuery

All images are published as multi-architecture OCI manifests supporting both linux/amd64 and Apple Silicon linux/arm64 natively.

Step-by-step guide

Mode 1: Single-container ad-hoc execution

You don't need to spin up a cluster or run any background services. The container's entrypoint acts as an intelligent capability dispatcher. You pass the tool name as the first argument, and it executes immediately:

1. Run the built-in self-test suite

Validates Spark DataFrame aggregations, Hive embedded metastore, Hadoop version, Java 21, and Python 3.12:

docker run --rm agentcloud/dataproc:3.0.0-debian13 self-test
2. Inspect component versions
# Check Spark 4.1.2
docker run --rm agentcloud/dataproc:3.0.0-debian13 spark --version

# Check Hadoop 3.5.0
docker run --rm agentcloud/dataproc:3.0.0-debian13 hadoop version

# Check Hive 4.2.0
docker run --rm agentcloud/dataproc:3.0.0-debian13 hive --version
3. Submit a PySpark workload

Mount your local script directory and run spark-submit directly inside the container. Here is a runnable example:

workloads/demo.py
from pyspark.sql import SparkSession
from pyspark.sql.functions import col, avg, count

spark = SparkSession.builder.appName("LocalDataprocDemo").getOrCreate()

data = [
    ("engineering", "Alice", 130000),
    ("engineering", "Bob", 115000),
    ("marketing", "Charlie", 95000),
    ("marketing", "Dana", 105000),
    ("data", "Eve", 140000),
]

df = spark.createDataFrame(data, ["dept", "employee", "salary"])

print("\n--- Local Dataproc PySpark Aggregation ---")
df.groupBy("dept").agg(count("employee").alias("headcount"), avg("salary").alias("avg_salary")).show()

spark.stop()

Execute it in one command:

docker run --rm \
  -v $(pwd)/workloads:/workload:ro \
  agentcloud/dataproc:3.0.0-debian13 \
  spark /workload/demo.py
4. Launch interactive PySpark or Spark SQL shells
# Interactive PySpark REPL
docker run -it --rm agentcloud/dataproc:3.0.0-debian13 pyspark

# Interactive Spark SQL CLI
docker run -it --rm agentcloud/dataproc:3.0.0-debian13 spark-sql
5. Standalone Hive queries & Bash access

The image ships with a schema-initialized embedded Derby metastore. You can execute HQL scripts instantly without spinning up an external database:

# Run an ad-hoc Hive query
docker run -it --rm agentcloud/dataproc:3.0.0-debian13 hive -e "SHOW DATABASES;"

# Drop into a raw bash prompt with all paths and Java environment configured
docker run -it --rm --entrypoint /bin/bash agentcloud/dataproc:3.0.0-debian13
Cluster mode

Mode 2: 3-Node Dataproc cluster with Docker Compose

When you need distributed execution — such as testing Spark-on-YARN job scheduling, verifying multi-node HDFS replication, or connecting to HiveServer2 over JDBC — the image includes an automated cluster supervisor.

docker-compose.yml
services:
  master:
    image: agentcloud/dataproc:3.0.0-debian13
    hostname: master
    container_name: dataproc-master
    environment:
      - CLUSTER_MODE=true
      - CLUSTER_ROLE=master
      - NAMENODE_HOSTNAME=master
      - RESOURCEMANAGER_HOSTNAME=master
      - METASTORE_HOSTNAME=master
      - WORKER_COUNT=2
      - HADOOP_CONF_DIR=/etc/hadoop/conf
    ports:
      - "19870:9870"    # NameNode Web UI
      - "18088:8088"    # YARN ResourceManager Web UI
      - "28080:18080"   # Spark History Server
      - "29888:19888"   # MapReduce Job History Server
      - "19083:9083"    # Hive Metastore Thrift
      - "20000:10000"   # HiveServer2 JDBC
    volumes:
      - master-hdfs:/var/hadoop/hdfs
      - master-metastore:/var/hive/metastore
      - master-warehouse:/var/hive/warehouse

  worker-1:
    image: agentcloud/dataproc:3.0.0-debian13
    hostname: worker-1
    environment:
      - CLUSTER_MODE=true
      - CLUSTER_ROLE=worker
      - NAMENODE_HOSTNAME=master
      - RESOURCEMANAGER_HOSTNAME=master
      - WORKER_COUNT=2
      - HADOOP_CONF_DIR=/etc/hadoop/conf
    volumes:
      - worker1-hdfs:/var/hadoop/hdfs
    depends_on:
      - master

  worker-2:
    image: agentcloud/dataproc:3.0.0-debian13
    hostname: worker-2
    environment:
      - CLUSTER_MODE=true
      - CLUSTER_ROLE=worker
      - NAMENODE_HOSTNAME=master
      - RESOURCEMANAGER_HOSTNAME=master
      - WORKER_COUNT=2
      - HADOOP_CONF_DIR=/etc/hadoop/conf
    volumes:
      - worker2-hdfs:/var/hadoop/hdfs
    depends_on:
      - master

volumes:
  master-hdfs:
  master-metastore:
  master-warehouse:
  worker1-hdfs:
  worker2-hdfs:

Cluster Service Endpoints

All services are forwarded to host ports designed to avoid common conflicts:

Service Host URL / Port Protocol
HDFS NameNode UI http://localhost:19870 HTTP Web UI
YARN ResourceManager UI http://localhost:18088 HTTP Web UI
Spark History Server http://localhost:28080 HTTP Web UI
Job History Server http://localhost:29888 HTTP Web UI
HiveServer2 jdbc:hive2://localhost:20000 Thrift / JDBC
Hive Metastore thrift://localhost:19083 Thrift Protocol
Cluster Commands
# 1. Start cluster in background
docker-compose up -d

# 2. Check NameNode and YARN nodes
docker exec dataproc-master hdfs dfsadmin -report
docker exec dataproc-master yarn node -list

# 3. Submit distributed job on YARN
docker exec dataproc-master spark-submit --master yarn /opt/localcloud/self-test/spark_smoke.py

# 4. Connect via Hive Beeline over JDBC
beeline -u "jdbc:hive2://localhost:20000" -e "SHOW DATABASES;"

# 5. Stop and tear down cluster (removes containers & volumes)
docker-compose down --volumes
Storage integration

Connecting to Cloud Storage & BigQuery

The container comes with Google Cloud Storage and BigQuery connectors pre-installed in the Spark classpath. You can connect in three ways:

Option A: Pure local storage (Offline)

Use standard POSIX local directories with -v $(pwd)/data:/data. Spark reads directly from file:///data/... without any cloud services.

Option B: Wire to LocalCloud emulators (Zero cloud cost)

If running alongside LocalCloud, provide emulator endpoint environment variables. The entrypoint automatically configures the GCS and BigQuery connectors:

docker run --rm \
  -e LOCALCLOUD_GCS_ENDPOINT=http://host.docker.internal:4443 \
  -e LOCALCLOUD_BIGQUERY_ENDPOINT=http://host.docker.internal:5388 \
  agentcloud/dataproc:3.0.0-debian13 \
  spark /workload/read_gcs_to_bq.py

Option C: Connect to real GCP with Application Default Credentials

When ready for staging validation against real GCS or BigQuery, mount your local ADC file:

docker run --rm \
  -v $HOME/.config/gcloud/application_default_credentials.json:/tmp/adc.json:ro \
  -e GOOGLE_APPLICATION_CREDENTIALS=/tmp/adc.json \
  -v $(pwd)/workloads:/workload:ro \
  agentcloud/dataproc:3.0.0-debian13 \
  spark /workload/real_gcs_pipeline.py
Automation

CI/CD automation: Running Spark tests in GitHub Actions

Because these images are 100% self-contained and run on public Docker runners, you can execute complete PySpark unit and integration tests directly in GitHub Actions without granting cloud permissions or creating GCP service accounts:

.github/workflows/spark-ci.yml
name: Spark Pipeline CI
on: [push, pull_request]

jobs:
  test-spark-pipelines:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Run PySpark Test in Dataproc 3.0 Container
        run: |
          docker run --rm \
            -v ${{ github.workspace }}:/workspace:ro \
            agentcloud/dataproc:3.0.0-debian13 \
            spark /workspace/tests/pipeline_smoke.py
Best practices

Performance and resource guidelines

Apple Silicon Native

Multi-arch linux/arm64 images run on M1/M2/M3/M4 Macs with zero emulation overhead, achieving near bare-metal execution speed.

Memory Allocation

Single-container jobs run comfortably with 2–4 GB memory. For the 3-node YARN cluster, configure Docker or Colima with at least 12 GB RAM.

Volume Persistence

Use docker-compose stop to pause the cluster and keep HDFS data intact, or pass --volumes on teardown for clean ephemeral resets.

Docker Hub repository

Available Dataproc image versions

All images are hosted on Docker Hub at hub.docker.com/repository/docker/agentcloud/dataproc. You can pin the exact Dataproc release your production workloads target:

Dataproc Version Docker Tag Spark Hadoop Hive Java Python
Dataproc 3.0 agentcloud/dataproc:3.0.0-debian13 4.1.2 3.5.0 4.2.0 Java 21 Python 3.12
Dataproc 2.3 agentcloud/dataproc:2.3.34-debian12 3.5.3 3.3.6 3.1.3 Java 11 Python 3.11
Dataproc 2.2 agentcloud/dataproc:2.2.85-debian12 3.3.2 3.3.6 3.1.3 Java 11 Python 3.10
Dataproc 2.1 agentcloud/dataproc:2.1.117-debian11 3.3.2 3.3.6 3.1.3 Java 11 Python 3.10
Dataproc 2.0 agentcloud/dataproc:2.0.161-debian10 3.1.3 3.2.4 3.1.3 Java 8 Python 3.8
Operational boundaries

Production validation boundaries

  • Component compatibility: These containers reproduce Apache Spark, Hadoop, and Hive runtime semantics. They do not simulate the Google Cloud Dataproc REST control plane (such as cluster provisioning RPCs, auto-scaling policies, or Compute Engine VM lifecycle).
  • Permitted use: The proprietary Public Preview License permits individuals and organizations, including for-profit companies, to use LocalCloud for non-production internal development, testing, CI, evaluation, and internal pilots.
  • Production verification: Always validate production jobs and infrastructure definitions against real Google Cloud Dataproc before deploying to production.