Free PROFESSIONAL-DATA-ENGINEER Practice Test Questions and Answers (2026)

View Mode
Q: 1
You are designing a data warehouse in BigQuery to analyze sales data for a telecommunication service provider. You need to create a data model for customers, products, and subscriptions All customers, products, and subscriptions can be updated monthly, but you must maintain a historical record of all dat a. You plan to use the visualization layer for current and historical reporting. You need to ensure that the data model is simple, easy-to-use. and cost-effective. What should you do?
Options
20 comments in the community discussion
1
B . Since the master node type is configurable in CUSTOM, I figured you can set numbers for all roles (masters, workers, parameter servers). Maybe a gotcha here about the master count but that's how I've always read it. Open to corrections if I missed something.
1
Option B, I thought you set counts for masters too, not just workers and parameter servers.
Q: 2
Which of these sources can you not load data into BigQuery from?
Options
31 comments in the community discussion
3
Option B
1
B vs A. If you just need the pipeline to use Shared VPC subnetworks, it's the service account that executes the pipeline requiring compute.networkUser, not the Dataflow service agent. Pretty sure that's what GCP docs say but always gets murky if org policies are custom. Disagree?
Q: 3
Your company is loading comma-separated values (CSV) files into Google BigQuery. The data is fully imported successfully; however, the imported data is not matching byte-to-byte to the source file. What is the most likely cause of this problem?
Options
29 comments in the community discussion
4
B . A is a trap since Kafka adds overhead, Pub/Sub is cheaper and survives unreliable leased lines.
1
B tbh, only edge case where A might win is if you needed local storage due to regulatory or privacy constraints. But question just wants cost-effective and resilient delivery, so managed Pub/Sub is perfect. Anyone disagree if on-prem cache wasn't a factor?
Q: 4
You work for an economic consulting firm that helps companies identify economic trends as they happen. As part of your analysis, you use Google BigQuery to correlate customer data with the average prices of the 100 most common goods sold, including bread, gasoline, milk, and others. The average prices of these goods are updated every 30 minutes. You want to make sure this data stays up to date so you can combine it with other data in BigQuery as cheaply as possible. What should you do?
Options
27 comments in the community discussion
1
It’s C for me. Not every Dataflow pipeline is a perfect directed graph, some complex ones can have variations or recursion. Data sharing feels possible via external storage, so I think C fits better. Disagree?
1
I don’t think it’s C. D is the trap here since pipelines aren’t supposed to share data between instances in Dataflow.
Q: 5
You launched a new gaming app almost three years ago. You have been uploading log files from the previous day to a separate Google BigQuery table with the table name format LOGS_yyyymmdd. You have been using table wildcard functions to generate daily and monthly reports for all time ranges. Recently, you discovered that some queries that cover long date ranges are exceeding the limit of 1,000 tables and failing. How can you resolve this issue?
Options
35 comments in the community discussion
5
Why not go with D? Only GCS lets your data survive after cluster deletion, not just restarts. Persistent disks (option B) are tied to the cluster lifecycle.
2
Its D
Q: 6
Your chemical company needs to manually check documentation for customer order. You use a pull subscription in Pub/Sub so that sales agents get details from the order. You must ensure that you do not process orders twice with different sales agents and that you do not add more complexity to this workflow. What should you do?
Options
29 comments in the community discussion
1
A and B are what I'd pick here. Scaling worker count or boosting instance type directly tackles CPU starvation in Dataflow. The buffer options (D, E) deal more with throughput spikes or persistent queueing, not pure compute limits. Pretty sure this is the intent, but open if anyone disagrees.
A and B tbh. Since CPU saturation is called out, just need to scale workers or use bigger instances. Adding buffers like in D/E isn't really the fix for compute issues.
Q: 7
To run a TensorFlow training job on your own computer using Cloud Machine Learning Engine, what would your command start with?
Options
32 comments in the community discussion
6
C . Had something like this in a mock and Google's guidance is to materialize dimensions using views when you need joins in a star schema, especially if you want to speed things up but not use more storage. Partitioning would help for date filters, but the question asks about storage impact specifically. Might be tr
4
C . Materializing dimensional data in views helps BigQuery run those star-schema joins more efficiently, and it doesn't bump up your storage bill. If the pain is with join performance and not just recent date scans, this fits. Anyone else think that's right?
Q: 8
Which methods can be used to reduce the number of rows processed by BigQuery?
Options
41 comments in the community discussion
1
Option D is the way to go. Pub/Sub plus Dataflow are both cloud-native, can autoscale automatically, and Dataflow supports windowed ordering for that 1-hour requirement. B is tempting if you like Kafka, but it isn't as integrated or fully managed in GCP as Pub/Sub. Anyone see a use case where B would actually be better
1
I don't think it's B. D is more cloud-native and actually autoscaling, plus Dataflow gives that windowed ordering you need.
Q: 9
You need to migrate a Redis database from an on-premises data center to a Memorystore for Redis instance. You want to follow Google-recommended practices and perform the migration for minimal cost. time, and effort. What should you do?
Options
32 comments in the community discussion
7
Option C here, as time travel is built-in for seven days and does not add extra costs. Clean and straightforward question.
4
Option C Had something like this in a mock and C was right for 7-day recovery, not D.
Q: 10
Which is the preferred method to use to avoid hotspotting in time series data in Bigtable?
Options
25 comments in the community discussion
4
Option B is correct. When your endpoint has an out-of-date SSL certificate, Cloud Pub/Sub can't confirm delivery due to failed TLS handshake, so it keeps retrying and you get duplicate messages. Option D is tempting since missing acks also trigger retries, but this scenario is specific to SSL issues with push endpoi
3
Option B actually makes sense here. If the SSL cert is out-of-date, Pub/Sub’s push will get handshake failures and treat it like a non-ack, which leads to retries and so you get duplicates. I know D's also common but in this scenario, cert problems will cause exactly this issue. Pretty sure B is right-correct me if I’m
Question 1 of 20

Premium Access Includes

  • Quiz Simulator
  • Exam Mode
  • Progress Tracking
  • Question Saving
  • Flash Cards
  • Drag & Drops
  • 3 Months Access
  • PDF Downloads
Get Premium Access
Scroll to Top

FLASH OFFER

Days
Hours
Minutes
Seconds

avail 10% DISCOUNT on YOUR PURCHASE