Why Engineers Are Choosing Simpler Databases Again

Why Engineers Are Choosing Simpler Databases Again
There's a quiet counter-revolution happening in how engineers are choosing data stores, and it doesn't have a snappy name or a CNCF working group. It's just people, quietly, reaching for Postgres or SQLite instead of whatever distributed thing was fashionable three years ago. Not because they've given up on scalability. Because they've done the maths.
The distributed systems tax is real, and a lot of teams have been paying it on workloads that didn't need to.
Why Distributed Systems Exist (A Brief Refresher You Probably Don't Need)
Start with the CAP theorem, because you can't really argue about databases without it. Brewer's conjecture — later proved by Gilbert and Lynch — states that a distributed system can provide at most two of three guarantees: Consistency, Availability, and Partition Tolerance. In practice, because networks partition (cables get cut, routers fall over, someone accidentally runs a while true loop in the wrong place), you're choosing between consistency and availability during a fault. CP systems sacrifice availability to stay consistent; AP systems keep responding but may serve stale data.
This leads naturally to the split between ACID and BASE. ACID (Atomicity, Consistency, Isolation, Durability) is what relational databases give you, and it's genuinely hard to implement across multiple nodes. A transaction that spans two machines requires coordination — typically two-phase commit or something equivalent — and coordination under partition conditions is exactly where the CAP theorem bites. The BASE model (Basically Available, Soft-state, Eventually consistent) relaxes those guarantees to keep the system humming through network hiccups. Your shopping cart might briefly show an item you've already deleted, but the system keeps taking requests. Which is fine until you're doing anything where "eventually" isn't acceptable.
To make distributed systems work in practice, you need a consensus mechanism: some way of agreeing which node is in charge. Raft is the current favourite for being comprehensible, though Paxos deserves credit for being the thing Raft was written to replace. The core idea is that the cluster elects a leader, all writes go through the leader, and the leader replicates changes to followers. If the leader disappears, the remaining nodes elect a new one — which requires a quorum, which is why you want an odd number of nodes and why "two-node cluster" is a phrase that should make you nervous.
Add sharding — Cassandra distributes its keyspace using consistent hashing so each node owns a segment of the ring and writes route accordingly; Redis Cluster does the same across 16,384 fixed hash slots — and you have a system that can sustain write throughput well beyond what a single machine can handle, with no single point of failure. Genuinely impressive engineering.
It's also a lot to operate.
The Tax
Every distributed system you run carries a fixed overhead. You need the consensus layer to maintain itself, which means your cluster needs at least three nodes to tolerate one failure, or five to tolerate two. You need monitoring that understands cluster topology, not just individual node health. You need runbooks for split-brain scenarios. You need someone who can read the logs when the leader election stalls at 3am.
Leadership transitions add latency spikes. Replication lag adds uncertainty about what "current" means. The operational surface is substantially larger than a single Postgres instance, and the failure modes are correspondingly more exotic. I have personally spent an afternoon debugging a Cassandra cluster that was technically healthy by every metric we were measuring, and technically not serving any reads. These things happen.
The question worth asking — and which not enough teams ask before reaching for Kafka or Cassandra — is whether the workload actually requires what distributed systems provide. High write throughput beyond what a single machine can sustain? Geographic replication across regions for low-latency global access? Genuine need to tolerate the loss of multiple nodes mid-operation? These are real requirements. They exist. They justify the tax.
But "our traffic might scale some day" is not a requirement. "The tech seems interesting" is not a requirement (though it is, at least, honest). "The last place I worked used it" is cargo cult, and cargo cult infrastructure is expensive both to build and to operate.
CAP Is Per Workload, Not Per System
One of the things that often gets lost in the CAP discussion is that you're not making a single choice for your whole architecture. You're making it, implicitly or explicitly, for each workload that needs persistence. When you run a single cluster and route everything through it, you've made a single CAP choice for everything — which is convenient, and frequently wrong.
At a previous role, we ran RabbitMQ as both a task queue and a message bus. Two jobs, sharing a broker. The problem was that the two types of data had fundamentally different requirements: real-time state tracking, where stale data was actively harmful, and durable task queues, where the message had to survive a partition even if that meant delayed delivery. State data should be optimised for availability; tasks should be optimised for consistency. Not subtle requirements. Not edge cases.
When we had a partition across two data centres, the same broker was trying to be correct about tasks and available for state simultaneously, and it couldn't be both. Eventually we recognised that we had conflated two problems. RabbitMQ's clustering model has tunable behaviour — you can configure queue durability, mirroring policy, and partition handling per queue — but reasoning about it correctly required making the workload boundaries explicit. We ended up running two separate clusters with different configurations, one for each data type. More infrastructure to operate, but radically simpler to reason about, and the failure modes became predictable in a way they hadn't been before.
The lesson wasn't "use two clusters." It was that the architecture had been obscuring a fundamental mismatch between workload requirements. Once you name the requirements clearly, the architecture follows — and sometimes the right answer is multiple simpler systems rather than one clever one.
The Simpler Path
The principle underlying the current drift toward simpler stores is one that software engineers are supposed to know and frequently forget when it comes to infrastructure: YAGNI. You Aren't Gonna Need It. Premature optimisation at the infrastructure layer is at least as costly as premature optimisation in code and considerably harder to roll back. You cannot git revert a Cassandra migration.
The reason for that is worth stating plainly, because it's not obvious until you've lived it. Migrating between data storage strategies isn't a configuration change — it's a data movement problem, and data has inertia. You have to read every record from the old store, transform it into whatever shape the new one expects, write it to the new store, verify it landed correctly, keep both systems synchronised while you cut over, and then unwind the old one. At any meaningful scale that process takes days to weeks of engineering time, runs alongside production traffic, and has a rollback story that's essentially "run the whole thing in reverse" — which is not nothing when the dataset has continued changing throughout. The bigger the dataset, the longer the window, the harder the cutback becomes. Unlike a bad code deployment, you can't simply take the previous container image and restart. The data is where it is.
The SQLite project's own documentation describes it as competing not with MySQL or Postgres, but with fopen(). It's a different problem: embedded, local, self-contained. A single file. No server process, no network round-trip, no connection pool. The latency floor for a query against a local SQLite database is measured in microseconds — not because SQLite is magic, but because it's in-process and the data is on the local filesystem, so the numbers are fundamentally different.
This doesn't mean it's only for toy projects. The SQLite project's own website runs on SQLite and handles something in the range of 400–500K HTTP requests per day. The maximum database size was raised a few years ago to 281 terabytes, which is further than most workloads will ever travel. The single-writer constraint is real — you can have many concurrent readers, but writes serialise — but for most read-heavy applications, and for any application where you can design around write concurrency, it's not a blocker.
The architecture that makes this work is one where the database is per-service, or even per-user. If each user's data lives in a separate SQLite file, you get effective sharding for free, and concurrent writes across users are no longer a concern because they're hitting different files. The server serialises requests per connection, and you can have thousands of connections each talking to their own database. This is not the architecture you'd reach for first if you'd spent the last decade thinking in terms of shared relational stores — but it's coherent, it's fast, and it's trivially operable.
Turso's libSQL extends this further. It's an open-source fork of SQLite that adds a server mode with HTTP access (useful for serverless environments that don't have a persistent filesystem), and local replica support — you can embed a full read replica in your application, syncing from a primary, which gets you microsecond-level read latency while still supporting a shared, writable source of truth. It's a reasonably elegant way to get the performance characteristics of embedded SQLite without giving up shared access, and it sidesteps the single-writer problem by pushing writes to a primary while reads stay local.
DuckDB sits at the other end of the same philosophy. Where SQLite is an OLTP store — row-oriented, optimised for transactional reads and writes on individual records — DuckDB is OLAP: columnar storage, vectorised execution, designed for analytical queries across large datasets. It also runs in-process, also requires no server, and can query Parquet, CSV, and JSON files directly without importing them. For workloads that previously would have reached for Redshift or BigQuery — aggregations, time series analysis, exploratory queries over flat files — DuckDB often does the job on a laptop, without a managed service contract. The two databases are not competing; they're complementary, covering different ends of the in-process spectrum.
Postgres as Kitchen Sink
Postgres has been quietly absorbing use cases that previously required specialist stores, and it deserves acknowledgement for this even if it makes the PostgreSQL ecosystem somewhat alarming to map.
The graph database case is instructive. Neo4j is a fine piece of software, and if your core model is genuinely a deeply connected graph — fraud detection, knowledge graphs, anything where relationship traversal is the primary operation — it earns its place. But a lot of workloads that reach for Neo4j are doing graph queries as one of several access patterns, not the only one. For those, Postgres with pgRouting does shortest-path, Dijkstra, and A* over standard graph data stored in regular tables. You don't get Cypher; you get SQL with graph-aware functions. Whether that's a limitation depends entirely on what you're building.
The vector story is similar. Purpose-built vector databases have proliferated over the last few years as embedding-based retrieval became mainstream, and several of them are excellent. But if you're already running Postgres and you need approximate nearest-neighbour search over embedding vectors, pgvector gives you that without adding another system to your stack. The scaling ceiling is lower than a dedicated solution, but for a lot of workloads — and especially for workloads that are already Postgres-shaped — it's the pragmatic choice.
Redis, Valkey, and the Question of Primitives
Redis occupies an interesting position in this discussion because it's neither a relational database nor a distributed system in the Cassandra sense, yet it's frequently reached for as a specialist store when its real strength is the richness of its primitives.
A Redis deployment gives you sorted sets, pub/sub, streams (Kafka-lite, for many workloads), geospatial indices, and — since Redis 8 — vector search. These are not things you'd build for yourself. Geospatial proximity queries, leaderboards, real-time pub/sub, embedding search: all available in a single process, operated with a single set of tooling. If your access pattern is well-served by those primitives, Redis is a very efficient choice.
Valkey is worth watching. It's the community fork that emerged after Redis changed its licence, and it's taking somewhat different positions — bloom filter support out of the box, scripting capabilities that are genuinely ahead of upstream Redis. It is developing its own character rather than just tracking Redis features, which is either exciting or worrying depending on how much you enjoy library bifurcation. Probably both.
The point here is not "use Redis for everything" — the persistence model requires care, and the memory-resident model has obvious limits. It's that when your workload fits the primitives, you can often avoid a purpose-built specialist store entirely by using Redis well.
A Few Rules of Thumb
Decision frameworks for data store selection get complicated quickly, so here are some that have held up in practice.
If you're starting a new service and don't yet have evidence of your load profile, Postgres is almost always the right default. It handles mixed workloads, it has the escape hatches (read replicas, logical replication, connection pooling), and the migration path to something more capable exists. Cassandra or CockroachDB on day one is a bet you're making with no data.
If your data is per-user or per-tenant and writes don't need to cross those boundaries, SQLite-per-tenant deserves serious consideration before anything else. The operational simplicity is worth more than the theoretical elegance of a shared relational store.
If your workload is analytical rather than transactional — you're aggregating over millions of rows rather than updating individual ones — DuckDB or a columnstore should be on your list before you reach for an OLAP managed service.
If you need a distributed store because the requirements genuinely demand it, separate your workloads by CAP profile first. Don't let a system that's optimised for availability serve data where staleness is harmful, and don't let a consistency-first system be the single point of failure for a service where degraded but available is acceptable. The RabbitMQ lesson scales.
If you're reaching for a specialist graph database, vector database, or time series store because of a single access pattern in an otherwise conventional application, check the extension ecosystem of whatever you're already running. The answer is quite often already there.
The AI-Assisted Tooling Gap
The shift happening in the industry isn't abandonment of distributed systems. It's a more careful interrogation of whether the requirements actually demand them. Part of what's driving that interrogation is AI-assisted development — and it's worth naming the double-edged quality of that specifically.
LLM tooling has materially lowered the barrier to reaching for an unfamiliar tool. You no longer have to default to the database you already know and can hand-write queries for; you can reach for something with better solution-fit to the actual problem, and get working code out the other end without spending a week reading documentation. That's a genuine improvement. Better tool-problem matching is good for systems design.
The gap it creates is operational. Knowing how to stand up DuckDB and write queries against it is not the same as knowing what happens when it falls over, what its failure modes look like, and how to recover from them. The experience that makes engineers reach for familiar tools isn't just familiarity — it's the accumulated scar tissue of having operated those tools under conditions that weren't in the docs. That scar tissue doesn't transfer automatically from an LLM. For new tools, you have to earn it, and the earning usually happens at an inconvenient time.
The practical implication: when you pick a tool because AI-assisted development made it accessible rather than because you have operational experience with it, build in explicit investigation of the failure modes before you depend on it in production. Not documentation reading — deliberate fault injection. Kill the process mid-write. Fill the disk. Partition the network. Find out what "broken" looks like before it happens without warning.
The Design Principle
The reason this isn't just "simple is good, distributed is bad" is that the constraints are real. If you need sub-100ms writes from users in Tokyo, Singapore, and São Paulo simultaneously, a single-region Postgres instance will fail you. If your write throughput exceeds what any single machine can sustain, SQLite won't save you. If your data model is fundamentally a graph and traversal depth is the performance-critical operation, you should probably talk to Neo4j before you spend three months fighting SQL.
Simpler stores are not universal. They're appropriate for the majority of workloads, which is not the same thing.
The pragmatic path is to design your architecture to play to the strengths of simpler stores and mitigate their weaknesses explicitly. Single-writer SQLite? Shard by user or tenant. Need reads from multiple regions? Run local replicas. Need more write headroom than Postgres can give you on a single box? First check whether you've exhausted what connection pooling, read replicas, and good indexing can deliver — because the answer is usually "not yet." When you genuinely hit the ceiling, move. Migrating to a more complex system when the requirements have proved it necessary is far less wasteful than running one speculatively for years.
There's a specific failure mode worth naming: the architecture that gets designed for the scale of a company in year five, built by a team of four, in year one. Distributed systems are not just operationally expensive — they impose coordination costs on development. Every engineer who joins the project needs to understand the consistency model, the replication topology, the failure modes. That knowledge transfer has a cost, and for early-stage teams it frequently exceeds the cost of the simplicity tradeoff they were trying to avoid.
The Actual Question
Complexity in infrastructure is not inherently bad. It exists because the problems it solves are real. The CAP theorem is not a historical curiosity; it's an operative constraint on every distributed system ever built. Election algorithms, consistent hashing, replication topologies — these are engineering achievements, and the systems built on them handle workloads that genuinely couldn't run any other way.
The question isn't whether distributed systems are good or bad. It's whether your workload justifies their cost. For write-heavy, globally distributed, high-availability applications with strong consistency requirements: almost certainly yes. For a SaaS application handling a few hundred requests per second, with a single-region deployment, where the bottleneck is almost certainly the application layer: probably not.
Distributed systems are a tax. Sometimes the tax funds something worth building. But running a Cassandra cluster for a hundred thousand rows and a few dozen requests per second is paying tax on income you haven't earned yet, in a currency that's harder to spend than it looks.
SQLite is on most of the devices you carry with you. It runs the Firefox profile store, the iOS contacts database, and roughly a third of the applications you used this week without noticing. The SQLite project's position is that it is the most widely deployed database engine in the world, and they're probably right. It is not a toy. It is what happens when a database is designed to solve the right problem rather than the impressive one.
That's the thing about "good enough" — it's only dismissive if you've misread what "enough" means.
References: SQLite Appropriate Uses · Turso/libSQL local replicas · Turso vector search · Supabase pgRouting · Neo4j · Redis 8 commands · Valkey