Professional Data Engineer

GCP_PDE
ProfessionalVersion PDEOfficial exam guide โ†—

Ready to test yourself?

A timed, blueprint-proportional exam drawn fresh from this bank โ€” with a per-domain score report.

๐Ÿ”’ Unlock the simulation โ†’

๐Ÿ”“ Free preview: showing 10 of 50 questions. Unlock the full bank โ€” every question, explanation, and reference.

Unlock all 50 questions โ†’

10 questions across 5 topics. Choose an answer for each question, then check it to see the correct answer and explanation.

Filter by topic

10 questions in Designing data processing systems

DESIGN_SYSTEMS

Designing data processing systems

10 questions in topic

Selecting storage and processing services, and designing batch/streaming pipelines and schemas.

1
Single choice~70s

Which service is the serverless data warehouse for large-scale SQL analytics on Google Cloud?

2
Single choice~90s

An IoT platform must store telemetry with millions of writes per second and low-latency operational reads. Which database fits BEST?

3
Single choice~90s

A system must ingest a continuous event stream, transform it, and make it queryable for near-real-time analytics. Which pipeline is the canonical design?

4
Single choice~90s

A team must migrate existing Apache Spark and Hive jobs to Google Cloud with minimal code changes. Which service fits BEST?

5
Single choice~90s

A new pipeline should be serverless, autoscaling, and handle batch and streaming with the same code. Which service should be chosen?

6
Single choice~110s

An application needs a relational database with strong consistency that scales horizontally across regions globally. Which service fits?

7
Single choice~90s

A team wants to build ETL pipelines visually (code-free) with prebuilt connectors, executing on Dataproc under the hood. Which service fits?

8
Single choice~90s

In a BigQuery-centric design, raw data is loaded first and transformed later with SQL in the warehouse. What is this approach called?

9
Single choice~100s

For analytical queries in BigQuery, which schema approach generally performs best?

10
Single choice~90s

Sales data is loaded once nightly for reporting and cost must be minimized. Which ingestion approach is most appropriate?