Pepitedata
  • Home
  • Services
  • Audits
  • Expert Call
  • About
  • Blog
  • Contact

object storage

A one kilobyte object fanning out into fifteen metadata copies, one per drive of the erasure set
object storage

Object Storage for Data Platforms at Scale

A production guide to running object storage at scale: why object count is a capacity axis, how prefix design and the background scanner set your real SLAs, and how to choose engines, media and release channels.

By Julien Laurenceau, 6 days2026-09-30 ago
Two industrial conveyor lanes in dark graphite, the upper lane continuous and the lower lane cut short with a small amber end cap
MinIO

MinIO Community in Production: Assess the Risk You Are Running

MinIO Community Edition has had no release since October 2025 and the repository is archived. Three CVEs will never be patched there. Here is how to assess what your cluster is actually exposed to, and where the maintained community fork fits in.

By Julien Laurenceau, 3 weeks2026-09-14 ago
Abstract illustration of multi-site object storage replication between data centers
MinIO

MinIO Site Replication: Sync Mode Will Not Save Your RPO

MinIO site replication has no cross-site quorum, and its replication queue exists only inside the source cluster: lose that site and the un-replicated backlog is gone and un-enumerable. Sync mode costs latency without fixing that. The precise model, the RTO/RPO table, and the DR patterns that match real SLAs.

By Julien Laurenceau, 2 months2026-08-03 ago
Two panels. Top: a single MinIO bucket overflowing with small files, a worried Tux, and a dead directory tree on a tombstone. Bottom: the same objects split across four prefix-partitioned buckets under separate folders, with a healthy green directory tree.
MinIO

MinIO on XFS: Inode Exhaustion and Prefix Design for Lots of Small Files

MinIO has no NameNode. It puts the namespace on XFS as real directories and xl.meta files. On Lots of Small Files that means you can exhaust inodes while df -h still looks fine, and a flat leaf can stall PUT, LIST, scanner, and ILM together. Here is the on-disk model, the inode math, and a prefix recipe that keeps XFS inside a regime you can operate.

By Julien Laurenceau, 2 months2026-07-27 ago
Diagram: MinIO erasure-coding storage inflation on small objects: per-drive xl.meta metadata replicated across the erasure set
data

Stop using MinIO as a NoSQL database: why S3 object stores collapse on small-file workloads

Two MinIO platforms, same root cause: used as a NoSQL store. Field notes on LIST IOPS, XFS directory limits, scanner & heal SLAs, the erasure-coding storage-efficiency inversion on small objects, and why Apache Cassandra (or Ceph) is the right answer on-prem in 2026.

By Julien Laurenceau, 5 months2026-05-17 ago
data

When the Lakehouse Works, but the Operating Model Breaks… Especially On-Premises

Note: This article was inspired by a LinkedIn post by Can Sinan A. on the hidden operational cost of an Iceberg migration, especially when governance moves from one access-control plane to several layers: catalog, storage and engine configuration Introduction Many organizations are modernizing their data platforms with lakehouse architectures, and Read more

By Julien Laurenceau, 5 months2026-05-04 ago
AI

Tesla’s .SMOL Format Shows Why Most Enterprise Data Lakes Are Architecturally Wrong

When Tesla published patent WO2024073080 describing a new file format internally called “.smol”, the headline was simple: 4x reduction in IOPS for AI training. Most people read this as a hardware story. It isn’t. It’s a data architecture story. And it exposes a structural weakness in how most enterprise data Read more

By Julien Laurenceau, 8 months2026-02-11 ago
ceph

Migrating Unmanaged OSDs to Managed OSDs in Ceph Squid

Migrating unmanaged OSDs is one of the operations I run for Ceph clients: Ceph storage consulting. If you are running Ceph Squid with a mix of managed and unmanaged OSDs, there is one thing worth stating clearly upfront: There is no in-place adoption mechanism in Squid. You cannot “import” an Read more

By Julien Laurenceau, 8 months2026-02-05 ago
apache spark

Apache Spark smoke tests on kubernetes

Introduction Benchmarking remains a critical (and often underestimated) tool when designing or validating large-scale data platforms. While many teams rely on synthetic workloads or production replays, standardized benchmarks still play a key role when comparing architectures, tuning clusters, or validating infrastructure choices. TPCx-HS is one of those benchmarks: designed to Read more

By Julien Laurenceau, 9 months2026-01-12 ago
data

The Hidden Costs of Using HDDs in On-Premises MinIO Deployments

Hidden storage costs are one of the targets of a data platform cost optimization review. Bring the IOPS dude ! As a solutions architect working with MinIO storage solutions, I’ve seen firsthand the challenges that come with on-premises deployments using hard disk drives (HDDs). While HDDs may seem like a Read more

By Julien Laurenceau, 10 months2025-12-15 ago

Posts pagination

1 2 Next
  • Privacy Policy
  • Mentions légales
  • CGV
  • Cookies
Hestia | Developed by ThemeIsle