Data Science

Your data lake is not the prerequisite

Why data science and AI platform projects fail on definitions, not on infrastructure, and how to let a real decision pull the data model into existence instead of waiting for the lake.

July 25, 2026
Share on
9 min read

There is a conversation I have had, in almost identical form, with a dozen clients. It happens early, usually in the second meeting, and it goes like this:

It is said with a certain relief, because it converts a hard organizational problem into a comfortable technical one. Finish the lake, then do the intelligent things. It is a clean sequence, easy to put on a roadmap, easy to fund.

It is also, in my experience, the wrong sequence, and more importantly, the wrong diagnosis.

I have seen companies with a modern lakehouse, immaculate ingestion pipelines, dbt models, a catalog, lineage, the whole apparatus, and a data science programme that produces nothing anyone acts on. I have also seen companies running on a handful of Postgres databases and a scheduled export, shipping a marketing mix model that reallocates seven figures of media spend, correctly, within a quarter.

About the author

Cyril Noirot

Cyril Noirot

Lead Data Scientist

Freelance data scientist. I design and ship decision systems — forecasting, pricing, marketing measurement, optimization.

Newsletter

Technical writing on forecasting, pricing, and decision systems. No fixed schedule, no spam.

Enter your email
Subscribe