Pepitedata
  • Home
  • Services
  • Audits
  • Expert Call
  • About
  • Blog
  • Contact

performance

Two bars: high throughput at 7 vCPU per item and collapsed throughput at 241 vCPU
data

Performance and Cost Engineering for Data Platforms

A field guide to performance and cost engineering: claim-checking with the scaling and queueing laws, measurement discipline, where GPU and CPU budgets leak, and how to turn a benchmark into a hardware decision.

By Julien Laurenceau, 5 days2026-09-30 ago
An open duct in a stone chamber holding a single file of evenly spaced glowing cubes, one entering and one leaving
data

Benchmark Best Practices, Part 2: Little’s Law Ties Throughput, Latency and Concurrency Together

Second in my benchmark best practices series. Little’s law checks that throughput, latency and concurrency describe the same system, exposes the load generator’s ceiling, and turns a latency budget into a sizing requirement.

By Julien Laurenceau, 2 weeks2026-09-22 ago
Two MinIO sites locked in a closed replication loop: delete markers stamped 405 circulate endlessly between two racks, piling up instead of converging
MinIO

MinIO site replication bug: how 684 million HTTP 405s exposed the delete-marker loop

HEAD probes on delete markers answer 405 and replication counts that as success. 20M stuck versions turned the loop into 684M useless requests in 48 h.

By Julien Laurenceau, 4 weeks2026-09-06 ago
A balance scale holding a small solid block on one pan and an oversized hollow sphere on the other
data

Benchmark Best Practices, Part 1: Amdahl’s Law Caps Your Scaling Claim

Part 1 of my benchmark best practices series: how I turn a scaling claim into an implied serial fraction with Amdahl’s law, and what the Universal Scalability Law adds. Includes a measured pipeline where assigning more vCPUs per image cut global throughput.

By Julien Laurenceau, 1 month2026-08-25 ago
Wireframe illustration of a LiDAR point cloud scan of powerlines and terrain, split into a dense grid of many small tile blocks on the left and a few larger tile blocks on the right, same data
object storage

Why File Count Matters as Much as File Size on Object Storage

Switching tile formats on a LiDAR ingestion pipeline cut output object count by roughly 25x at the same data volume, independent of compression. Here is why object count is its own cost and performance axis on object storage, and what to check before you scale a pipeline that writes many small files.

By Julien Laurenceau, 1 month2026-08-22 ago
Isometric illustration of a distributed edge computing network of glowing server and sensor nodes, with one node shaped like a retro video game controller, on a dark navy and cyan background
data

Which Message Broker for IoT and Real-Time Control

Choose a message broker for real-time IoT control: start from the latency budget, test ack mode and packet loss, and when to skip the broker for UDP.

By Julien Laurenceau, 1 month2026-08-22 ago
Abstract illustration of multi-site object storage replication between data centers
MinIO

MinIO Site Replication: Sync Mode Will Not Save Your RPO

MinIO site replication has no cross-site quorum, and its replication queue exists only inside the source cluster: lose that site and the un-replicated backlog is gone and un-enumerable. Sync mode costs latency without fixing that. The precise model, the RTO/RPO table, and the DR patterns that match real SLAs.

By Julien Laurenceau, 2 months2026-08-03 ago
Two panels. Top: a single MinIO bucket overflowing with small files, a worried Tux, and a dead directory tree on a tombstone. Bottom: the same objects split across four prefix-partitioned buckets under separate folders, with a healthy green directory tree.
MinIO

MinIO on XFS: Inode Exhaustion and Prefix Design for Lots of Small Files

MinIO has no NameNode. It puts the namespace on XFS as real directories and xl.meta files. On Lots of Small Files that means you can exhaust inodes while df -h still looks fine, and a flat leaf can stall PUT, LIST, scanner, and ILM together. Here is the on-disk model, the inode math, and a prefix recipe that keeps XFS inside a regime you can operate.

By Julien Laurenceau, 2 months2026-07-27 ago
Erasure coding is efficient on big files but degrades to many inefficient copies on small files
MinIO

MinIO and Small Files: When Erasure Coding Becomes 15x Replication

MinIO fixed the HDFS NameNode limit, but it has no index: it writes one xl.meta per object on every drive of the erasure set. On 5 servers of 12 NVMe, MinIO picks a 15-wide set by default, so each small object is stored 15 times over. A worked sizing that looks fine for two years and dies in days.

By Julien Laurenceau, 2 months2026-07-23 ago
Abstract visualization of Sail engine bridging Rust and Spark technologies
apache spark

Sail: When Apache Spark Meets Rust (A Practitioner’s Deep Dive)

Sail is an open-source Apache Spark replacement written in Rust. It drops the JVM, speaks Spark Connect, and runs 4 to 6 times faster than Spark with native accelerators on ClickBench. A deep dive into its architecture, benchmarks, and production readiness.

By Julien Laurenceau, 4 months2026-06-14 ago

Posts pagination

1 2 Next
  • Privacy Policy
  • Mentions légales
  • CGV
  • Cookies
Hestia | Developed by ThemeIsle