Ceph in Production: RBD, RGW and CephFS
A field guide to Ceph in production: which surface serves which workload, what failure domains cost in replicas, how the upgrade lifecycle really rolls back, and what key rotation does not do.
A field guide to Ceph in production: which surface serves which workload, what failure domains cost in replicas, how the upgrade lifecycle really rolls back, and what key rotation does not do.
Ceph tenants and accounts solve different problems. Separate the Tentacle upgrade from account adoption, then test ownership, IAM and notifications.
Switching tile formats on a LiDAR ingestion pipeline cut output object count by roughly 25x at the same data volume, independent of compression. Here is why object count is its own cost and performance axis on object storage, and what to check before you scale a pipeline that writes many small files.
MinIO site replication has no cross-site quorum, and its replication queue exists only inside the source cluster: lose that site and the un-replicated backlog is gone and un-enumerable. Sync mode costs latency without fixing that. The precise model, the RTO/RPO table, and the DR patterns that match real SLAs.
MinIO has no NameNode. It puts the namespace on XFS as real directories and xl.meta files. On Lots of Small Files that means you can exhaust inodes while df -h still looks fine, and a flat leaf can stall PUT, LIST, scanner, and ILM together. Here is the on-disk model, the inode math, and a prefix recipe that keeps XFS inside a regime you can operate.
MinIO was not built for Lots of Small Files (LOSF). No global index, no read repair, and a scanner that can take weeks to notice silent corruption. Here is what breaks, why tiering makes it worse, and when you should use a database instead.
Two MinIO platforms, same root cause: used as a NoSQL store. Field notes on LIST IOPS, XFS directory limits, scanner & heal SLAs, the erasure-coding storage-efficiency inversion on small objects, and why Apache Cassandra (or Ceph) is the right answer on-prem in 2026.
Update, September 2026: the sibling failure hit site replication. ILM expiration on a replicated mesh created delete markers that MinIO re-probes forever: 706 million replication requests in 48 hours, 96.8 percent answered 405, no data moved. Full write-up: how 684 million HTTP 405s exposed the delete-marker replication loop. Introduction In Read more