Independent data platform audit for performance, reliability and technical due diligence.
Independent technical review before an investment, acquisition, platform commitment or business-critical scaling decision.
Independent expert review for investors, buyers, CTOs, Heads of Platform and engineering teams evaluating Spark, Kafka, Kubernetes, object storage, Ceph/MinIO, Cassandra, private cloud or AI/data infrastructure.
I help decision-makers identify the real technical risks, bottlenecks, reliability gaps and cost drivers in complex data systems – before capital, roadmap or vendor commitments become hard to reverse.
When to call me
You should call me when:
- You are investing in, acquiring or buying into a company whose value depends on a technical platform
- You are about to subscribe to a strategic data, AI or infrastructure product and need an independent technical read
- Spark jobs are too slow, unstable or expensive
- Kafka pipelines are lossy, slow or hard to operate at scale
- Your data platform is blocking AI or analytics initiatives
- Kubernetes / object storage / private cloud migration is becoming messy
- Ceph, MinIO, Cassandra or storage layers create latency or reliability issues
- Infrastructure costs are rising and nobody can explain why
- You need to size a new platform and choose the right servers before committing budget
- Your disaster recovery plan has never been challenged, and the stated RTO / RPO must survive a real failover
- Your team needs an independent architecture review before a costly decision
What I deliver
Expert Call – €250 HT / 1h
A focused video session to unblock one architecture, performance, scalability or cost issue.
Best for: validating a technical direction; challenging an architecture proposal; identifying next actions on a painful issue; getting an external read before a larger decision.
Technical Due Diligence – From €4,500 HT
An independent technical audit for investors, buyers and leadership teams before investing in a company, acquiring shares, validating a vendor platform or subscribing to a strategic data/AI infrastructure product.
Typical scope: architecture and codebase review where available; infrastructure, scalability and reliability risk assessment; technical debt and operational maturity review; vendor or target-company claims challenged against evidence; written risk memo with go/no-go considerations and prioritized follow-up questions.
Best for: investors who need technical diligence before a transaction, executives evaluating a platform-dependent acquisition, or teams committing to a product where performance, lock-in, reliability or cost could materially affect the business case.
Flash Architecture Review – From €4,500 HT
A focused review of one critical technical decision, one bottleneck or one system component.
Typical scope: architecture review; technical interviews; metrics, logs or configuration review when available; written diagnosis; prioritized recommendations; restitution call with CTO / tech lead / platform team.
Disaster recovery is a frequent focus. I review DR architecture documents before implementation to catch the anti-patterns that would make the stated RTO / RPO unreachable. After implementation, I verify that the DR document describes what actually runs. The target on paper must survive a real failover.
Platform Sizing – From €4,500 HT
A focused engagement to size or right-size a data infrastructure platform – for a new deployment, a capacity uplift, a migration, or a re-platforming decision.
This is one of my strongest areas. I size open-source components – Kafka, Cassandra, Ceph, PostgreSQL, Kubernetes, Spark, Pulsar – and translate the result into concrete hardware: the right instance types from a public cloud catalog, or the right server configuration from an OEM catalog such as Dell or HPE.
Typical scope: target workload analysis; sizing of compute, storage and network for Kafka, Cassandra/ScyllaDB, Ceph/MinIO, PostgreSQL, Kubernetes, Spark or Pulsar clusters; server and instance selection from public cloud or OEM catalogs (Dell, HPE); capacity headroom recommendations; cost estimate against the chosen infrastructure (private cloud / on-prem / public cloud); written sizing document and restitution call.
Data Platform Performance Audit – From €7,500 HT
Typical range: €7,500-€15,000 HT depending on scope, urgency and system complexity.
A deeper audit for complex data, cloud or AI platforms where performance, reliability or scalability issues have business impact.
Typical deliverables: architecture map; bottleneck analysis; reliability and scalability risks; infrastructure cost and capacity insights; quick wins; prioritized roadmap; executive restitution for CTO / VP Engineering / Head of Platform.
Fractional Performance Advisor – From €3,000 HT / month
Ongoing senior guidance for teams operating complex data infrastructure.
Best for teams that need a few hours per month of independent senior review, architecture challenge, incident analysis or roadmap support – without hiring a full-time principal engineer.
Who this is for / who this isn’t for
This is for: investors, acquirers, CTOs, Heads of Platform, VP Engineering, Heads of Data and engineering teams evaluating or operating production data, cloud or AI infrastructure where performance, reliability, scalability, technical debt or cost has measurable business impact.
This is not for: teams that need staff augmentation, single-contributor freelance work, generic cloud strategy, or projects under €5k of scope. If you need fewer than 5 days of focused work, the Expert Call is the right entry point.
Systems I review
The audit focuses on the systems that usually create hidden cost, reliability risk and performance drag in mature data platforms.
Spark on Kubernetes
I review executor sizing, CPU overcommit, shuffle behavior, autoscaling assumptions, queueing, namespace limits, pod failures, object-store access patterns and job timelines. The goal is to separate Spark tuning problems from Kubernetes scheduling, storage or network bottlenecks.
Object storage, Ceph and MinIO
I look at latency, throughput, S3 operation rates, erasure coding, disk layout, network paths, metadata pressure, small-file behavior, lifecycle rules and failure domains. Object storage often looks healthy at the bucket level while Spark, AI training or analytics workloads suffer from request amplification and unstable tail latency.
Cassandra and ScyllaDB
I review data models, partition size, compaction, tombstones, JVM and GC behavior, repair practices, node balance, latency percentiles and incident history. For operational databases, peak throughput matters less than predictable latency. The important question is whether the platform can keep predictable latency under real production load.
Kafka, private cloud and AI infrastructure
I examine broker sizing, retention, consumer lag, network topology, storage contention, GPU utilization, batch scheduling and capacity allocation. In private-cloud environments, the same physical bottleneck can affect Spark, Kafka, object storage and AI workloads at the same time.
Signals I examine
A useful audit starts from evidence. I typically review job timelines, Prometheus and Grafana dashboards, CPU and memory profiles, JVM metrics, S3 latency and throughput, I/O patterns, network paths, pod restart history, scheduler events, queueing, Cassandra compaction statistics, GC pauses, incident postmortems and cloud or private-cloud cost allocation.
I also review the team’s own hypotheses. Many platform teams already know part of the answer, but they lack time, external distance or cross-system evidence. The audit turns those signals into a clear diagnosis and a prioritized set of decisions.
Deliverables
- Architecture map: the systems, dependencies, data paths and ownership boundaries that matter for performance and reliability.
- Bottleneck analysis: the constraints that explain slow jobs, unstable latency, rising cost or reliability incidents.
- Risk register: reliability, scalability, operability and vendor-risk items that leadership should track.
- Sizing and cost model: compute, storage, network and capacity assumptions with cost impact where data is available.
- 30/60/90-day roadmap: prioritized fixes, quick wins, decisions to postpone and changes that need deeper engineering work.
- Executive summary: a clear version for CTO, VP Engineering, Head of Platform, investor or buyer discussions.
What this is not
This is not staff augmentation, a generic cloud migration project, vendor implementation, a digital transformation program or a long freelance mission. The work is diagnostic and decision-oriented. It helps you decide what to fix, what to avoid and where engineering time has the highest return.
What I need from you
The best inputs are architecture diagrams, metrics dashboards, configuration excerpts, incident reports, workload samples, cost reports, capacity plans and interviews with the engineers who operate the platform. Production access is often unnecessary. Good evidence and direct technical conversations are usually enough to find the main constraints.
Selected client work
A non-exhaustive list of past engagements. I name some clients with their permission and anonymize the others.
- BPCE-IT – Cassandra performance and reliability work supporting an API Management platform. “His expertise and skills have helped improve the performance and behavior of our platforms.”
- Multi-petabyte private cloud data lake – independent audit of a large on-premise data lake supporting analytics and AI workloads
- MinIO and Ceph at scale – performance and observability work on production object storage
- Spark on Kubernetes – performance and reliability engineering on private-cloud Spark platforms
- Geospatial / satellite imagery pipelines – data infrastructure design for high-volume image processing
References available on request under NDA.
Why me
I combine PhD-level numerical optimization, HPC background and 15+ years of hands-on production experience in distributed systems, data platforms and private cloud.
I have worked on multi-petabyte data lakes, Spark on Kubernetes, Ceph/MinIO object storage, Cassandra/ScyllaDB, satellite imagery pipelines, banking infrastructure, telecom and aerospace systems.
I do not sell generic cloud advice. I help teams make better technical decisions when systems become too slow, too costly, too fragile or too complex to reason about internally.
Typical outcomes
Depending on context, clients use the audit to:
- Reduce processing time
- Lower infrastructure cost
- Identify storage or compute bottlenecks
- De-risk a migration
- Improve production reliability
- Challenge vendor, target-company or integrator recommendations
- Surface technical risks before an investment, acquisition or strategic platform subscription
- Prioritize engineering work
- Avoid over-engineering and unnecessary cloud spend
Common audit scenarios
Most teams do not call for an audit because one dashboard looks bad. They call when several systems interact badly and the internal diagnosis is no longer obvious. These are the patterns I see most often.
Spark jobs are slow, unstable or expensive
A Spark problem can come from code, partitions, shuffle, file layout, Kubernetes limits, object storage, autoscaling, queueing or a mix of all of them. I review job timelines, stage metrics, executor behavior, pod scheduling, storage latency and data layout together. This avoids tuning one layer while the real constraint sits somewhere else.
The output is not a generic list of Spark best practices. It is a ranked diagnosis: which jobs create the cost, which bottlenecks are structural, which fixes are quick, and which changes require engineering work or platform decisions.
Object storage works, but analytics workloads suffer
Ceph, MinIO and S3-compatible platforms can look healthy from the storage dashboard while data workloads see unstable performance. Small files, request amplification, erasure coding, disk contention, network paths and metadata pressure can all create a gap between raw storage health and application performance.
I connect storage metrics to workload behavior. That includes S3 operation rates, tail latency, object size distribution, retry patterns, throughput by client, network placement and failure domains. The goal is to identify whether the platform needs layout changes, capacity changes, operational changes or application-side changes.
Cassandra or ScyllaDB latency becomes unpredictable
Latency incidents in Cassandra-like systems often come from a combination of data model drift, tombstones, compaction, repairs, node imbalance, JVM pressure, disk behavior and client access patterns. Average latency rarely tells the story. The important signals are percentiles, distribution, hot partitions, compaction backlog, read amplification and incident timing.
The audit separates operational hygiene from architectural risk. Some issues need configuration or repair-process changes. Others point to a data model that no longer matches production traffic. Leadership needs to know which case they face before assigning months of work.
Infrastructure cost rises without a clear owner
Cost problems in data platforms are rarely owned by one team. Spark sizing, Kubernetes quotas, storage layout, replication, retention, GPU allocation, idle capacity and vendor pricing all contribute. I review cost drivers with the technical evidence behind them, so the team can distinguish waste from necessary headroom.
The result is a cost model tied to engineering choices. That may show an immediate right-sizing opportunity, a storage-layout problem, a workload scheduling issue or a platform decision that needs executive attention.
A strategic decision needs independent technical due diligence
Investors, buyers and leadership teams often need a fast technical read before a platform commitment, acquisition, vendor selection or major roadmap decision. I focus on the evidence that changes the decision: scalability claims, operational maturity, hidden platform risk, dependency risk, cost exposure and the credibility of proposed fixes.
The deliverable is written for decision-makers, but grounded in technical review. It gives leadership a clear view of what is healthy, what is risky, what needs follow-up and which claims should be challenged before money or roadmap time is committed.
How I separate symptoms from root causes
I start by mapping the data path, not by assuming the problem sits in the component with the loudest alert. A slow Spark job may expose object storage latency. An object storage incident may come from network placement. A Cassandra latency spike may come from compaction pressure created by a product change. A GPU cost issue may come from pipeline scheduling, not from GPUs.
I look for cross-system evidence: timestamps that line up, workload changes, capacity limits, retries, queuing, tail latency, saturation, noisy neighbors and operational events. This makes the audit useful even when the team already has strong engineers. The value is the independent synthesis across platform layers.
The final recommendations are ranked by business impact, implementation effort and risk. Some fixes are immediate. Some require a design decision. Some should be rejected because they add complexity without enough return. A good audit should help you spend less time debating symptoms and more time making the right technical decisions.
What leaders can decide after the audit
The audit is built to support concrete decisions, not to produce a long report that nobody owns. Typical decisions include whether to resize Spark workloads, change object-storage layout, adjust Kubernetes quotas, postpone a migration, challenge a vendor claim, split a platform roadmap, or fund a reliability workstream before adding new product scope.
For investors and buyers, the output highlights the technical risks that can affect valuation, integration cost, delivery timelines or operating margin. For CTOs and platform leaders, it clarifies where engineering time should go first and which ideas are expensive distractions.
How the audit is scoped
A good scope is narrow enough to produce a clear answer and broad enough to avoid local optimization. I usually start with the business question: a slow reporting chain, unstable AI workload, expensive Spark platform, unreliable storage layer, Cassandra latency problem, private-cloud sizing decision or technical due diligence before a strategic commitment.
From there, I define the systems in scope, the evidence available, the people to interview, the decision deadline and the expected output. If the useful answer fits in an Expert Call or a Flash Review, I will say so. If the risk is cross-system, the full audit gives enough time to connect workload behavior, infrastructure metrics, operational history and cost impact.
This keeps the engagement efficient. The goal is not to inspect every dashboard. The goal is to find the evidence that changes the decision.
How it works
- Scope – Clarify the issue, system boundaries, urgency and expected business impact.
- Diagnose – Review architecture, metrics, configuration, incidents, team hypotheses and constraints.
- Prioritize – Identify bottlenecks, risks, quick wins and high-ROI fixes.
- Restitution – Deliver a clear written diagnosis and a prioritized technical roadmap.
Related technical notes
These notes show the kind of production issues an audit can surface and prioritize:
- Spark on Kubernetes CPU overcommit and operator patterns
- MinIO performance with large numbers of small files (LOSF)
- GPU infrastructure cost and overprovisioning strategies
- Data lakehouse operating models at scale
FAQ
Do you need production access?
Usually no. Metrics, architecture diagrams, configuration excerpts and technical interviews are often enough to start.
Can this be done remotely?
Yes. Most audits are remote-first.
Is this only for large companies?
No, but the issue must be costly enough to justify a premium audit.
Can you help implement the recommendations?
Yes, but the audit is designed to stand alone. Implementation can be scoped separately.
Why not just hire a freelance for a few months?
Because you may not need more hands. You may need an independent diagnosis before spending months in the wrong direction.
Ready to scope an audit?
A free 15-min call to discuss your platform, your bottleneck or which service fits, and decide next steps. That could be Technical Due Diligence, the Audit, a Flash Review or an Expert Call.