Ceph in Production: RBD, RGW and CephFS
A field guide to Ceph in production: which surface serves which workload, what failure domains cost in replicas, how the upgrade lifecycle really rolls back, and what key rotation does not do.
A field guide to Ceph in production: which surface serves which workload, what failure domains cost in replicas, how the upgrade lifecycle really rolls back, and what key rotation does not do.
Ceph tenants and accounts solve different problems. Separate the Tentacle upgrade from account adoption, then test ownership, IAM and notifications.
Ceph 20.2.4 and 19.2.6 patch four authentication and authorization CVEs. The remediation also needs CephX key rotation, client checks, and RGW multisite planning.
A practical Ceph RGW upgrade plan for multi-site S3 platforms, covering replication validation, canary rollout, S3 acceptance tests and the limits of rollback.
Migrating unmanaged OSDs is one of the operations I run for Ceph clients: Ceph storage consulting. If you are running Ceph Squid with a mix of managed and unmanaged OSDs, there is one thing worth stating clearly upfront: There is no in-place adoption mechanism in Squid. You cannot “import” an Read more
Introduction Benchmarking remains a critical (and often underestimated) tool when designing or validating large-scale data platforms. While many teams rely on synthetic workloads or production replays, standardized benchmarks still play a key role when comparing architectures, tuning clusters, or validating infrastructure choices. TPCx-HS is one of those benchmarks: designed to Read more
Choosing a PVC backend is an architecture decision I review; the Ceph storage consulting page describes the cluster side. Most discussions about PVC focus on on-prem deployments, but many of these technologies are equally relevant in the cloud, especially when using managed block storage like AWS EBS, which caps at Read more