Pepitedata
  • Home
  • Audits
  • Expert Call
  • About
  • Blog
  • Contact

data

Two identical machine units labeled PUBLIC TIER and VETTED TIER under a bracket reading SAME WEIGHTS, the left one with its output slot bolted shut and marked REFUSED over a bare floor, the right one with the same slot open and output pouring into a heap
AI

After Coldcard, AI Found 1,029 Bugs. Censored Models Found None.

A volunteer red team filed 4,962 AI findings across 390 repositories in 27.5 hours, and not one came from a public US frontier model. Not a capability gap: the capable versions sit behind identity gates. Refusal policy is now where the labs differentiate, and it belongs on your dependency list.

By Julien Laurenceau, 2 weeks2026-08-07 ago
A vast monolithic structure of glowing strata sending a single beam of light across darkness into a small object that glows from within
AI

How Big Models Teach Small Models, and Why It Looks Like School

Knowledge distillation is a large model teaching a small one, and it works almost exactly like school: soft labels instead of a bare answer key, pairing weeks instead of reports, and a tutor who marks your own attempt rather than a textbook of last year’s solutions. The limits map too, and the last one costs money when you size the hardware.

By Julien Laurenceau, 2 weeks2026-08-06 ago
A lone cyclist ahead in turbulent airflow while a tight bunch of riders follows in smooth sheltered air
AI

The Frontier Model Has No Teacher. That Is Why Leading Costs Ten Times More.

A frontier lab optimizes an absolute target and has to explore. A challenger optimizes a distance to the leader and can exploit. That is the whole cost asymmetry, and it shows up in GPU hours and in a bike race.

By Julien Laurenceau, 2 weeks2026-08-06 ago
Audio waveform processed by a local GPU into structured transcript files
AI

Transcribe 100 Hours of Podcasts with whisper.cpp

I used whisper.cpp on an RTX 5070 to transcribe about 100 hours of podcasts in roughly three hours, then turned the Markdown output into a searchable prompt knowledge base.

By Julien Laurenceau, 2 weeks2026-08-04 ago
Abstract illustration of multi-site object storage replication between data centers
MinIO

MinIO Site Replication: Sync Mode Will Not Save Your RPO

MinIO site replication has no cross-site quorum, and its replication queue exists only inside the source cluster: lose that site and the un-replicated backlog is gone and un-enumerable. Sync mode costs latency without fixing that. The precise model, the RTO/RPO table, and the DR patterns that match real SLAs.

By Julien Laurenceau, 2 weeks ago
Abstract dark navy graphic: local model nodes linked to a sober data-plane silhouette
AI

Local LLMs, Political Bias, and Why llama.cpp Still Matters

Public frontier models lean progressive on open benchmarks. Local inference via llama.cpp is still the practical way to keep an uncensored second opinion. Memory is the bottleneck; TurboQuant-style KV compression is what will make local writing workers routine in multi-model agents.

By Julien Laurenceau, 2 weeks2026-08-02 ago
Two panels. Top: a single MinIO bucket overflowing with small files, a worried Tux, and a dead directory tree on a tombstone. Bottom: the same objects split across four prefix-partitioned buckets under separate folders, with a healthy green directory tree.
MinIO

MinIO on XFS: Inode Exhaustion and Prefix Design for Lots of Small Files

MinIO has no NameNode. It puts the namespace on XFS as real directories and xl.meta files. On Lots of Small Files that means you can exhaust inodes while df -h still looks fine, and a flat leaf can stall PUT, LIST, scanner, and ILM together. Here is the on-disk model, the inode math, and a prefix recipe that keeps XFS inside a regime you can operate.

By Julien Laurenceau, 3 weeks2026-07-27 ago
Erasure coding is efficient on big files but degrades to many inefficient copies on small files
MinIO

MinIO and Small Files: When Erasure Coding Becomes 15x Replication

MinIO fixed the HDFS NameNode limit, but it has no index: it writes one xl.meta per object on every drive of the erasure set. On 5 servers of 12 NVMe, MinIO picks a 15-wide set by default, so each small object is stored 15 times over. A worked sizing that looks fine for two years and dies in days.

By Julien Laurenceau, 4 weeks2026-07-23 ago
Abstract network of three connected data-center clusters representing a Ceph multi-site S3 upgrade
ceph

Upgrading a Ceph Multi-Site S3 Platform from Squid to Tentacle

A practical Ceph RGW upgrade plan for multi-site S3 platforms, covering replication validation, canary rollout, S3 acceptance tests and the limits of rollback.

By Julien Laurenceau, 1 month2026-07-17 ago
AI

Model Fusion Beats the Frontier Model. The Cost Case.

OpenRouter put numbers on a pattern agent builders already use: a panel of models with a judge synthesizer beats the best single model, and a budget panel matched frontier quality at half the cost. What that means for your AI platform.

By Julien Laurenceau, 2 months2026-06-23 ago

Posts pagination

1 2 3 Next
  • Privacy Policy
  • Mentions légales
  • CGV
  • Cookies
Hestia | Developed by ThemeIsle