Google Cloud Architect: Data Platform Selection
Data-platform questions on the Google Professional Cloud Architect exam are rarely solved by memorizing product names. The current guide asks architects to choose data-processing solutions, storage types, database approaches, data-transfer configurations, retention and lifecycle, data growth, security, latency, and protection. The decision begins with workload behavior and operational requirements, then maps those requirements to a service.
Object, file, and cloud data models differ in access pattern, protection model, consistency, and operating responsibility across providers. At the PCA level, candidates need to go further: identify access pattern, consistency needs, transaction semantics, query shape, scale, regional requirements, recovery objectives, data lifecycle, integration, and team operating model before selecting the platform.
The best architecture often uses more than one data service. The goal is not to consolidate everything into a single “strategic database,” but to keep the number of platforms small enough to operate while still matching materially different workload needs.
Start with the data model and operations: key-value, document, relational, time-series, analytical, object/blob, file, streaming event, or graph-like relationship. Then describe read/write ratio, transaction size, concurrency, query complexity, latency, consistency, retention, and expected growth. Those characteristics usually eliminate more choices than a feature checklist.
Also identify who operates the data and how frequently the schema or workload changes. A team that needs strong relational transactions has a different problem from an analytics pipeline that scans petabytes, even if both systems store “customer data.”
Operational systems optimize for serving application requests with predictable latency and transaction semantics. Analytical systems optimize for scanning, aggregating, and joining large data sets. Trying to make one platform excel at both patterns can create cost and performance problems.
An architect should define how operational data reaches the analytical environment, how fresh it must be, and whether replication, change data capture, batch pipelines, or streaming is justified. This also creates a security boundary that needs explicit access and retention rules.
Object storage works well for durable unstructured data, artifacts, backups, data lakes, and content accessed through object APIs. File storage is appropriate when applications need shared filesystem semantics. Those models differ in access pattern, namespace behavior, performance, and migration compatibility.
Avoid choosing file storage only because an existing application used a traditional NAS. Determine whether the application truly depends on filesystem semantics or whether modernization can simplify operations by moving to object or managed database access.
Relational workloads need to be separated by scale, availability, transaction, compatibility, and geographic requirements. Some applications fit a managed relational service with familiar SQL semantics. Others need horizontally scalable, globally distributed transactions and can justify the operational and modeling consequences of that choice.
The exam rewards matching the data requirement rather than selecting the most powerful service. If a regional application can meet its objectives with a simpler managed database, a global platform may add cost and design complexity without business value.
Event-driven architectures need decisions about ingestion rate, ordering, replay, retention, processing windows, late data, and consumer independence. The data platform must support the way downstream systems recover from failures and reprocess events.
Do not treat a messaging or streaming service as a database substitute. Determine where durable system-of-record data lives, where transient events live, and which processing layer creates derived data. Clear ownership makes recovery and governance easier.
The PCA guide explicitly includes data security, encryption, secrets, customer-managed keys, regulation, ownership, and data sovereignty. A technically ideal platform can be wrong if it cannot meet residency, audit, or key-management requirements in the required region or operating model.
Identity design matters as well. Cloud identity and access separates least privilege, service identities, and administrative authority. For data architecture, map application identity, administrative identity, encryption-key access, and analytical access separately instead of granting one broad data role.
Storage allocation, lifecycle, retention, growth, backup, and recovery are all explicit exam considerations. Forecast volume and access pattern rather than only current size. Data that grows tenfold can change partitioning, cost, index strategy, archival policy, and recovery time.
Migration planning should also include transfer bandwidth, cutover, validation, rollback, schema conversion, dual-write or synchronization periods, and downstream consumers. The safest migration is often staged so correctness can be proven before the legacy system is retired.
For exam practice, use the PCA scenario-based preparation approach: state the requirement, list the decisive constraints, choose the platform category and service, then explain which trade-off you accepted. If the requirement changes—global writes, lower RPO, different residency, new analytical scale—explain whether the choice still holds.
This produces architecture reasoning that survives product updates. Services evolve, but data-access patterns, transaction semantics, failure models, security boundaries, and cost-performance trade-offs remain the durable decision framework the Professional Cloud Architect role requires.
Data-platform sprawl is an operational cost even when every individual choice is technically defensible. Each service adds IAM patterns, monitoring, backup behavior, cost models, expertise requirements, and incident procedures. Prefer a smaller portfolio when requirements are close enough, and introduce a new data technology when the difference in workload need is significant enough to justify the operating burden.
Query isolation and workload management can be as important as raw scale. If operational and analytical users compete for the same resources, one workload can degrade another even when the database is technically capable of both. Separate systems or resource controls can provide more predictable service levels, but they add data movement and governance overhead. The decision should reflect which form of isolation the business actually requires.
Data locality influences both performance and cost. Moving large data sets between regions, clouds, or on-premises systems can dominate latency and egress expense. Place processing near the data when practical, and design migration or replication flows intentionally. A compute service that is cheaper per unit can still be the wrong choice if it requires constant movement of high-volume data.
Schema evolution should be part of platform selection. Some applications require strict relational contracts and carefully controlled migrations; others benefit from flexible document shapes or append-oriented event models. Flexibility shifts work rather than eliminating it: downstream consumers still need rules for compatibility, validation, and interpretation when fields change.
Operational tooling matters during incidents. Ask how the platform exposes slow queries, replication lag, hot partitions, quota pressure, storage growth, failed jobs, or backup status. Teams should be able to identify whether a problem is application logic, query design, service capacity, network path, or platform state without depending on opaque trial and error.
Finally, document the exit path. Managed services provide significant operational value, but enterprise architecture should know how data can be exported, transformed, retained, or deleted if the workload changes. This is not an argument against managed platforms; it is a way to make the accepted portability trade-off visible when the service is selected.
Backup and restore characteristics should be compared before a platform is approved. Understand the available backup types, retention limits, point-in-time recovery behavior, cross-region options, restore time at expected data scale, and how encryption keys participate in recovery. A nominal RPO is not enough if the restore procedure cannot meet the RTO.
Data governance can also change the topology. Separate raw, curated, and serving data when ownership or retention differs; apply labels and metadata that support discovery and cost attribution; and define deletion or legal-hold behavior. The platform should make the required governance practical rather than depending on every consumer to remember policy manually.
Performance testing should use representative data volume and query shape. Small development data sets hide partitioning, indexing, skew, network transfer, and concurrency problems. Before committing to a service, test the workload characteristics that drive cost and latency, especially for analytical scans, bursty ingestion, or globally distributed access.
Cost models should include the way a service is actually consumed. Provisioned capacity, per-operation pricing, storage tiers, data processing, replication, backups, network transfer, and idle development copies can shift the economics substantially. Estimate representative steady-state and peak workloads, then identify which parameter will dominate cost as the system grows. A platform that is inexpensive at pilot scale may become the wrong choice under production access patterns.
Data quality and lineage requirements can also influence the architecture. Analytical and AI workloads often need to know where data came from, which transformations were applied, and whether a field is trustworthy enough for a decision. Select platforms and pipeline patterns that make those controls practical. If lineage or quality evidence must be reconstructed manually after an incident or audit, the data architecture is operationally incomplete.
Platform ownership should be explicit before production. Define who tunes performance, approves schema changes, manages backups, monitors cost, rotates credentials, and handles incidents. A technically suitable database without an operating owner becomes fragile because essential maintenance is assumed rather than assigned.
Assign ownership explicitly.
A final architecture review should also ask how the choice will be reversed. Migration cost, export paths, schema portability, retention, and operational ownership matter because a platform decision is stronger when the organization can evolve it without an emergency redesign.
