Skip to main content
Postgres, DynamoDB, Redis, Elasticsearch, Snowflake. We use all five for different workloads. The decision criteria, not the marketing comparison.

Cloud-Native Databases: Choosing the Right Database for Your Workload

KU
Kiril Urbonas
10 months ago • 7 min read•Updated 2 days ago•59 views

Postgres, DynamoDB, Redis, Elasticsearch, Snowflake. We use all five for different workloads. The decision criteria, not the marketing comparison.

Key takeaways

  • A cloud-native database is one built to run on cloud or Kubernetes infrastructure: it scales out or automatically, replicates across zones, fails over without a human, and is run as a managed service or operator.
  • Default to managed Postgres, add DynamoDB for high-throughput key-value access, and add Redis, search, or a warehouse only when a workload demands it.

A cloud-native database is a database designed to run on cloud or Kubernetes infrastructure: it scales horizontally or automatically, replicates across availability zones, fails over without a human, and is consumed as a managed service or operator. Our verdict after years of running five of them: default to managed Postgres, and add others only for access patterns Postgres can't serve.

What makes a database cloud-native#

The CNCF defines cloud native as building and running workloads "at scale in a programmatic and repeatable manner" on public, private, or hybrid cloud. For a database that translates into four properties:

  • Elastic scale: capacity grows by adding nodes or by the provider autoscaling, not by buying a bigger box.
  • Built-in replication and failover: copies live in more than one zone and a primary loss is handled automatically.
  • Declarative operation: clusters, backups, and upgrades are API calls or Kubernetes resources, not runbooks.
  • Usage-shaped cost: you pay for what runs, and some options scale to zero.

Examples span both models: managed services like Aurora, DynamoDB, and Snowflake, and open-source projects like Vitess (sharded MySQL, a CNCF Graduated project since November 2019) and CloudNativePG (a Kubernetes operator for highly available Postgres). We run roughly five in production: Postgres on RDS, DynamoDB, Redis on ElastiCache, OpenSearch, and Snowflake. This is the decision tree we use when a new service asks "what should I store this in?"

The default: Postgres#

Most workloads start in Postgres unless there's a specific reason otherwise. The relational model is understood by nearly every developer; transactions, foreign keys, and joins work as expected; backups, failover, and monitoring are mature; pgvector adds vector search and JSONB handles semi-structured data. We run about 30 Postgres databases, mostly RDS Multi-AZ, holding user and transactional data.

Postgres stops being the answer when the access pattern is a single-key lookup at very high throughput (DynamoDB), the data is hot and ephemeral (Redis), the query is full-text search (Elasticsearch or OpenSearch), or the work is large read-only aggregation (Snowflake or BigQuery).

DynamoDB: key-value at scale#

DynamoDB is right for very high write throughput, predictable single-key lookups, single-digit-millisecond latency, and workloads that need to scale out without a migration. We use it for session state, idempotency keys, high-volume event streams with TTL expiry, and feature-flag state.

It is bad at joins, aggregations, and range queries on non-key fields. It performs by your access pattern, so a new query shape means a redesign. Reads are eventually consistent by default; strongly consistent reads are opt-in per request. Done well, with single-table design, GSIs for alternate access patterns, and sort keys that encode hierarchy, it scales beautifully. Treated like a relational database, it gets expensive and slow.

Redis: ephemeral hot data#

Redis is right for caching, sorted-set leaderboards, counters, rate limiting with token buckets, lightweight pub/sub, and distributed locks with caveats. It is fast partly because durability is optional, so it is wrong for data you can't lose, and RAM makes 100 GB-plus datasets expensive. We never use Redis as a primary store; there is always a source of truth in Postgres or DynamoDB that can rebuild the cache.

Elasticsearch and Snowflake: search and analytics#

Elasticsearch (or OpenSearch) handles full-text search, faceted filtering, and aggregations over logs. We use it for knowledge-base search, the hot tier of our logging stack, and product search. It is eventually consistent, has no transactions, and is heavy to operate; we moved from self-managed to AWS OpenSearch and the premium was worth it.

Snowflake is the OLAP side of the OLTP/OLAP split: product analytics, financial reporting, cohort analysis, and joins across Postgres, Salesforce, and Stripe data replicated in with Fivetran. Sub-second operational queries and per-record CRUD are not what it is for.

The decision tree#

  1. Transactional, relational, under a few TB? Postgres.
  2. Key-value, very high throughput? DynamoDB.
  3. Ephemeral hot data with a backing source? Redis.
  4. Full-text search over documents? Elasticsearch or OpenSearch.
  5. Analytics over large data? Snowflake.
  6. Vectors for ML retrieval? pgvector if you're on Postgres, a dedicated store only past where that stops scaling; see choosing a vector database.
  7. High-volume time series? ClickHouse or Timescale.

Most workloads end at step 1 or 2.

Common antipatterns#

Postgres for high-volume event ingestion. A service writing 50k events per second to Postgres will have a bad time; use DynamoDB or Kafka plus aggregation.

DynamoDB as a relational database. Six GSIs later, costs explode and consistency gets weird. Postgres was the right answer.

Redis as a primary store. Fine until Redis goes down and the data is gone.

Snowflake behind a user-facing dashboard. Queries take 5-10 seconds; cache aggregates in Postgres or DynamoDB instead.

Operational comparison#

DBManaged costOps timeRestore time
Postgres (RDS Multi-AZ)$$LowHours (point-in-time)
DynamoDB$$$Very lowMinutes (PITR)
Redis (ElastiCache)$$LowFast (rebuild from source)
Elasticsearch (managed)$$$MediumSlow (snapshot restore)
Snowflake$$$$Very lowTime Travel built in

We have migrated in both directions: Postgres to DynamoDB for a session store struggling at 30k writes per second, DynamoDB back to Postgres when a feature's queries grew complex, and analytics from Postgres to Snowflake. Each took weeks per service.

Questions people ask#

What is a cloud-native database?#

It is a database built for elastic, failure-prone cloud infrastructure rather than a single server. It scales out or autoscales, keeps replicas in several availability zones, fails over automatically, and is operated through APIs or Kubernetes resources. Aurora, DynamoDB, CockroachDB, and Snowflake are managed examples; Vitess and CloudNativePG bring the same properties to MySQL and Postgres on your own Kubernetes clusters.

Is Postgres a cloud-native database?#

Postgres itself is a single-primary server, but run as a managed service like RDS Multi-AZ or Aurora, or through a Kubernetes operator like CloudNativePG, it gets the cloud-native properties: automated failover, declarative backups, and replicas across zones. What it doesn't get is write scale-out. When one primary is the bottleneck, look at DynamoDB, Vitess, or CockroachDB.

Which cloud-native database should a new team start with?#

Managed Postgres plus Redis for caching covers most products. Add DynamoDB when a clear, stable key-value access pattern outgrows Postgres, search when users need full-text queries, and a warehouse when analytics slow down production. For spiky traffic, scale-to-zero options remove idle cost; we compared them in best serverless databases.

The call we'd make#

Default to managed Postgres; it fits most workloads and the operations are mature. Use DynamoDB when the access pattern is clear and stable, Redis only with a source of truth, and search engines and warehouses only for search and analytics. Most teams don't need five databases; we ended up there because each workload had a different shape. If read latency across regions is the driver rather than data model, that is a different question, covered in edge databases for low-latency apps.

Sources#

Explore topics:Cloud
React

Get the DevOps Troubleshooting Cheat Sheet

Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.

Share this post
KU

Kiril Urbonas

AI Engineer

560 articles
View all articles by Kiril Urbonas

You might have missed

Evergreen posts worth revisiting.