Three lanes: a time series waveform, a record grid and ledger rows each aligned to its matching engine slot

Choosing a data store by product name is how platforms end up with a time-series engine holding transactional records, or a general database paying rent as a metrics dump. The choice is made by three questions plus one: what shape the data has, how it is accessed, what consistency the workload actually requires, and what the engine’s operational defaults will do to your history over time. This guide maps the engine classes against those questions and collects what the store posts on this blog established: one silent data loss story from lifecycle defaults, and one performance story from specialization. It also sets the slots for the cluster’s planned spokes on Cassandra and time-series platforms beyond the metrics stack.

The one-line version: data shape, access pattern and consistency choose the class; operational defaults decide whether the class stays safe.

Why store selection goes wrong

Two patterns recur. Teams pick by popularity or by a benchmark headline, then meet the failure mode of the class they bought: a lifecycle engine quietly deleting history on data that cannot be regenerated, or a general distributed design losing badly to a constrained one on the exact workload shape the business runs. Both failures are predictable before purchase. Neither is visible on the first deploy day.

The selection frame: data shape, access pattern, consistency

Data shape first: append-mostly measurements keyed by timestamp are time series; keyed entities with updates and occasional queries are records; account movements with balances that must never drift are ledger entries. Access pattern second: range scans over time, key-direct reads and writes, or high-contention transactional writes. Consistency third: eventual or tunable for most operational data, serializable where money or double-entry correctness is non-negotiable, and a written answer to what a successful acknowledgment actually promises. The fourth criterion is the hidden one: operational defaults (retention, compaction, cleanup, replication) are a data-loss contract, and both case studies in this guide turn on it.

Time-series engines: what a TSDB is actually for

A time-series database is built for append-mostly metrics that age out through retention tiers and resolution tiers (raw, then downsampled). That lifecycle machinery is the product: it is what makes “infinite retention” affordable, and it is where the risk moves. Prometheus sets the pattern for short-term storage; systems like Thanos extend it onto object storage for long-term history. The lifecycle’s assumptions are the contract: if downsampling must produce a replacement before raw blocks expire, then a silent downsampling failure becomes permanent data loss the moment retention runs. The worked failure is in the section below. A second time-series platform with a different lifecycle and query model (Warp10) is a planned spoke of this cluster, as are the wider TSDB families beyond the metrics stack.

Three bars of equal height labelled raw, 5m and 1h joined by an equality marker
The retention contract in one picture: raw, downsampled and twice-downsampled durations set equal.

Distributed databases: records at scale and the consistency dial

Wide-column and distributed SQL engines answer the keyed-records shape at scale: high write throughput, multi-node or multi-DC replication, and tunable consistency per query. The design axis is what happens on write acknowledgment and on conflict: quorum systems price latency for consistency where you ask for it, and last-write-wins styles can silently discard the correct write when clocks drift. Cassandra and its peers are covered in depth in a dedicated guide planned for this cluster; until it lands, the engagement page for Apache Cassandra and NoSQL data modeling covers the review scope. The selection rule here: replication topology is not a consistency guarantee, and “distributed” is a topology, not a promise.

Purpose-built engines: when a specialized ledger wins

The third class trades breadth for guarantees: one fixed workload, one data model, and correctness properties a general engine offers only with effort. The archetype is a double-entry financial ledger where balances must never drift and storage faults are assumed, not hoped against. The engineering answer is specialization all the way down: a single-core write path that removes contention instead of managing it, memory fixed at compile time so nothing grows under load, and hardware-aligned records. The measured case (including a passed serializability audit and a corruption-tolerance story) is in the TigerBeetle deep dive. The selection rule: purpose-built wins when the workload is fixed and the failure mode of the general engine is exactly the risk you are buying the new store to remove.

Two failure stories: silent data loss and the wrong access pattern

The first story is lifecycle-driven. In a Thanos deployment, the compactor merges blocks, produces downsampled resolutions and enforces retention. If downsampling reports success while producing empty output (it can, with a warning), retention still deletes the raw blocks and no replacement exists. With retention durations misaligned (say raw kept a year and downsampled six months), the configuration is exactly the one where history disappears quietly. The defensive rule is an equality, not a preference: raw, 5-minute and 1-hour retention set to the same duration, compactor warnings under alert, and a staging rehearsal of compaction failure. The full analysis is in how Thanos defaults can lead to silent data loss.

Two panels: a small console card with a skip series warning and a large empty frame labelled history deleted
The warning was small. The outcome was a year of history.

The second story is shape-driven: a general distributed database asked to run a high-contention financial workload, paying coordination costs a constrained ledger engine would never incur. Both stories share one lesson this blog keeps re-meeting in storage, replication and now stores: quiet failure is the default mode of infrastructure, and the defaults are the contract. The same theme runs through replication drops into unreadable counters and deletion loops counted as success.

The store-selection checklist

Four answers before any benchmark. Data shape: time series, keyed records, documents or ledger entries. Access pattern: time-range scans, key-direct reads, high-contention writes or secondary queries. Consistency: eventual, tunable or serializable, and what write acknowledgment must mean. Operational defaults: retention, compaction, replication and cleanup behavior, read as a data-loss contract and verified before trusted. Only then do benchmarks and hardware sizing become meaningful, and the vendor’s cost or speed claims get tested on your own workload shape. When a store decision is about to set your reliability and cost for years, capacity planning and platform sizing or a data infrastructure audit is where I pressure-test it.

FAQ

How do I choose between a time-series database and a general-purpose database?

On data shape and access pattern. If the records are append-mostly measurements keyed by timestamp, read mostly as time ranges, and age out through retention and resolution tiers, a time-series engine is built for exactly that lifecycle. If the records are keyed entities with updates, joins or transactions, a general-purpose database fits better. Forcing one class to do the other’s job makes a TSDB’s retention machinery a data-loss risk on irreplaceable data, and makes a general database an expensive metrics dump.

What is the most common way teams silently lose historical metrics data?

Misconfigured lifecycle settings, not query bugs. When retention deletes raw blocks and downsampling can succeed while producing no output, history disappears with no error reaching an operator. The defensive rule: align every retention duration to the same value, alert on compactor warnings such as empty chunk skips, and simulate compaction and downsampling failures in staging. Treat retention and compaction defaults as a data-loss contract.

When is a purpose-built database worth adopting?

When the workload is fixed, the correctness requirement is non-negotiable, and the general engine’s failure modes are precisely the risk the new store is meant to remove. A double-entry financial ledger is the archetype: one narrow data model, extreme write contention, and a need for strong serializability under storage faults. For everything else, the specialization cost usually outweighs the gain.

Does choosing a distributed database automatically improve reliability?

No. Replication, quorum and consistency semantics differ per engine and per query, and misreading them is its own failure mode: asynchronous mirroring is not consensus, and last-write-wins conflict resolution can discard the correct write when clocks drift. Reliability comes from matching the engine’s consistency model to what the workload requires and verifying the defaults that implement it.

Related posts


0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *