Skip to content
Vevano

Why a data platform

Why decisions get better when the data has one home.

Most companies do not lack data. They lack one place where the data can be trusted — and therefore something to decide on. Which customers come back? Where does the margin disappear? What happens if we cut the price?

A data platform is the shared foundation underneath: where data lands once, is described once and is used by many. It stops being a reporting system someone visits at month end and becomes the place the company is steered from: finance, operations, product and sales read the same numbers, and what happens in one part of the business is visible in another while it happens. The same number holds for the board and for the person with the case in front of them.

That is what becoming data-driven means: questions can be asked and answered while they still matter. The behaviour that comes before a customer leaves is visible while there is still time to act on it; the cost that always follows a particular kind of order can be pointed at, not just suspected. Anyone can buy the tools — your history is your own, and that is where the edge is.

And it is the prerequisite for AI. The foundation comes before the model, not after: a model cannot be better than the data it builds on, and without one place with unambiguous data, clear access and history, the answer is something nobody can explain or defend. Data first, model after — that order is not up for debate.

one foundation, two readers

01

Ready for AI

An agent looks for data the way a person does. If it cannot find them, nobody else can either.

Most AI in production is not about the model. It is about the data: where it comes from, who is allowed to use it, and whether the answer can be traced back.

An agent looks for data the way a person does: it has to find the right table, understand what the fields mean, and trust that the number holds. If you cannot find and trust the data in your own company, no agent can either.

A model that searches your own data inherits its quality. If it finds three copies of the same customer with three addresses, it answers with one of them — and nobody can say which. The platform makes the data unambiguous before the question is asked.

Access is the other half. An interface where someone can ask their way to something they are not allowed to see is not a shortcut. It is a leak. The answer should follow the same rules a report does, and there should be a trail showing which fields were used.

It is also what makes an answer worth trusting: when the data has provenance and history, what the model proposed can be checked against what was actually there.

from many connections to one source

02

From silos to one source

Why the question "which number is right?" becomes a meeting, and what changes that.

It starts innocently. One connection between two systems is cheap, quick and entirely right. The trouble is that the connections grow faster than the systems: every new source that has to reach every report becomes another copy that can go stale.

The copy gets its own definitions. "Active customer" means one thing in the CRM, something else in the accounting system, and a third thing in the report someone built last year. "Which number is right?" is no longer a lookup. It is a meeting — and in that meeting, whoever speaks loudest wins.

A platform turns the connections around: sources connect to the platform instead of to each other. One landing place, one definition per concept, and a line back to the source for every number shown.

Not everything should move. Data that a single system uses is better off where it is. The platform gathers what has to be used across systems — and leaves the rest alone.

files, a log, any reader

03

Open table formats

Delta Lake and Apache Iceberg: the files sit in your own storage.

Object storage made storage boring: almost unlimited data, cheap per gigabyte. What was missing was the table.

An open table format gives it back: the table is Parquet files with a transaction log, in your own storage. Object storage then gets what only databases had: transactions, an enforced schema, versions, and points in time to go back to.

Two open formats have become the two names to know: Delta Lake and Apache Iceberg. Both build on the same principles and are read by a wide range of engines, both are developed in the open — and they converge year by year. Apache Hudi is the third in the same family. The category therefore matters more than the individual tool: an open format is a choice that does not lock in the rest.

The practical part is that the reader does not decide. A write becomes whole or not at all, a reader sees a consistent picture while data is being written, and every write is a version with a timestamp. History and streaming can live in the same table, and every reader sees the same truth — with Delta or with Iceberg.

That is why the format matters more than it sounds. The data is files, in your storage, in a publicly described format. They open with a tool other than the one the platform uses, and they move without anyone having to be asked first. The day a vendor relationship ends, the data does not end with it.

owner, access, provenance, quality

04

Governance and trust

Who owns the data, who may see it, and why do we trust it?

Governance often arrives with the smell of paperwork. It is really four questions the platform has to answer in seconds: Who owns this? Who may see it? Where did it come from? What has happened to it?

The first answer is a catalogue: one place where datasets are described and searchable, generated from the platform rather than from a spreadsheet kept beside it. The second is an owner per dataset. A person, not a committee — and someone who actually knows what the data means.

The third is access, set per dataset, column and row, with masking where someone needs to see that a person exists without seeing who it is. The rule should be simple: access is granted because someone is going to use the data, not because nobody said no.

The fourth is provenance: which sources and which steps produced the number, and who read or changed it, and when. Quality belongs in the same place — expectations about a dataset, such as required fields, correct format and unique keys, tested and reported. A break should be found by the system, not by the accounts.

Two years later, "which fields went into that decision?" is a lookup, not a story. A rule nobody can enforce is just an intention — which is why governance belongs in the platform, not in a document beside it.

now, and every point behind it

05

Now and then, in the same table

History and real time are not two projects when the format allows both.

History and real time sound like two solutions with two budgets. With a format that keeps versions, they are two sides of the same table: the latest moment for the person watching, and any earlier point for the person explaining.

That gives a property that is easy to underestimate. A report can be read as it actually was. If someone corrects a number today, both versions can be brought up side by side — and then "why did it change?" is answered, not explained away.

Real time is less about speed than about distance: how far behind is the number? Some decisions tolerate a day, others tolerate minutes. What matters is that the choice is made per stream and written down — not that everything has to be instantaneous.

one copy, and movable

06

Cost and ownership

One copy instead of many, and a cost you can read off your usage.

A platform costs twice: for storage and processing, and on the day something goes wrong. The second bill is the larger one, and cheap tools do not make it go away.

One copy instead of five makes the arithmetic readable. You pay for the storage you actually use, and processing can be priced per question. No metering per seat or per lookup — no bill that grows because more people started using it.

What matters most is still the contract on the day someone wants to move on. Open formats and standards mean the data can be read without the original tool, and the platform can be installed somewhere else without being rebuilt. Locked data is more expensive than expensive tools.

We are not saying this is cheaper than anything else. We are saying the cost is something you can read off your own usage — and that it does not grow on its own.

The test

Six questions you can ask

  • 01

    Can we read the data without the vendor?

    Ask for the format, not for reassurance. The files should open with a tool other than the one the platform uses.

  • 02

    Can we move it?

    A platform that can be installed somewhere else without being rebuilt is the only one you can leave.

  • 03

    Do we know who has seen what?

    Access and audit trail per dataset, not per system.

  • 04

    Can we explain a number?

    From the source, through the steps, to the report or the model. Today, not as a story afterwards.

  • 05

    What happens if someone leaves?

    Documentation, code and access should be yours. A platform that only works with one person present is a risk.

  • 06

    Can we say no to a dataset?

    Governance that cannot refuse access is not governance. It should be possible to keep something out.

Next step

Talk to us

Tell us briefly what you need. We will reply with how we can help — or with who can help better.