Skip to content

Data Platform & Analytics

One governed place where your data lands, is modelled, and means the same thing to everyone.

What this is

A data platform is the layer between your source systems and every report, model or AI assistant that reads from them. It ingests batch and streaming data, lands it, transforms it into modelled tables, and exposes a semantic layer where a business term has exactly one definition. Catalogue, lineage and access control are part of the platform, not an addition to it.

The failure mode is not technical. It is three departments presenting three different revenue figures in the same meeting because each built its own extract with its own filters. Tooling does not fix that; a definition with a named owner does. We also see platforms migrated onto a modern warehouse with the old ungoverned habits carried across intact.

When you need it

If more than one of these is true, this is usually the right place to start.

  • Two teams bring different numbers for the same metric and the meeting turns into a reconciliation exercise.
  • Analysts spend most of their time assembling extracts rather than answering questions.
  • Nobody can trace a figure in a report back to the source system it came from.
  • An AI or reporting initiative has stalled because the underlying data is not trusted or not reachable.

What the scope covers

  • Source inventory and ingestion design across batch, streaming and change data capture.
  • Warehouse or lakehouse architecture sized to your query patterns rather than to a vendor reference diagram.
  • Transformation layer with tested, version-controlled models and business logic written down.
  • Semantic layer so each metric has one definition, one owner and one implementation.
  • Governance: data catalogue, lineage, access control and a data quality testing regime.

What you receive

DeliverableWhat it contains
Source and metric inventoryEvery source system, the metrics drawn from it, and the conflicting definitions currently in circulation across teams.
Ingestion and storage layerBatch, streaming and change data capture pipelines landing into a warehouse or lakehouse with clearly defined zones.
Modelled and semantic layerVersion-controlled transformations and a metric layer in which each definition carries a named owner.
Governance and quality controlsCatalogue entries, lineage, role-based access and automated tests that fail a pipeline before bad data reaches consumers.

Reference architecture

A reference, not a template. Your estate decides which parts apply and in what order they arrive.

Data platform and analytics reference architecture: ingest, platform and governance layersIngest: Batch & Streaming, Change Data Capture, Landing Zone. Platform: Warehouse / Lakehouse, Transformation, Semantic Layer. Governance: Data Catalogue, Lineage, Access ControlIngestBatch & StreamingChange Data CaptureLanding ZonePlatformWarehouse / LakehouseTransformationSemantic LayerGovernanceData CatalogueLineageAccess Control
Data platform and analytics reference architecture: ingest, platform and governance layers

How success is measured

Targets are agreed with you before the work starts, and reported against for its duration.

  • Data freshness against the agreed service level for each pipeline, measured from source event to availability.
  • Pipeline test coverage, and the proportion of runs stopped by a quality check before data reaches consumers.
  • Share of published metrics carrying a named owner, a written definition and traceable lineage.

Questions we are asked

  • Warehouse or lakehouse?

    It depends on your workloads, not on which is newer. Structured reporting over relational sources is well served by a warehouse; unstructured data, large files and machine learning feature pipelines argue for a lakehouse. Many organisations end up with both, and the decision that matters is where the governed, modelled tables live.

  • Do we need to replace our BI tool?

    Usually not. Most reporting problems sit upstream of the BI tool - undefined metrics, untested transformations, unclear ownership. Fixing the modelling and semantic layers often makes the existing tool perfectly acceptable. We would rather change one layer than run a tool migration that solves nothing.

  • How does this relate to our AI plans?

    Directly. Retrieval, agents and predictive models all read from this layer and inherit both its quality and its permissions. Building AI on ungoverned data produces confident answers derived from the wrong numbers, which is worse than no answer at all.

  • Where does the data live, given residency rules?

    That is a design input, not an afterthought. Regional cloud regions, in-country deployment and hybrid patterns are all workable, and each carries cost and operational consequences we set out before you choose. Sector regulation in your market usually narrows the options quickly.

  • How do we stop the old habits coming back?

    By making the governed path the easy path. If the modelled tables are complete, documented and fast, private extracts stop being worth the effort. Where they persist it is usually a gap in the model rather than a discipline problem, and the fix is to close the gap.

  • How long does this take?

    The first governed domain - one business area, its sources and its metrics - is a matter of weeks and should deliver something usable. A platform covering the whole organisation is a programme, and it should be sequenced by business value rather than by finishing the ingestion of every source first.

Continue reading

  • AI & Data

    The full domain, and the other capabilities within it.

  • AIOps

    Correlation, anomaly detection and noise reduction across metrics, logs and traces, with runbook automation kept under explicit human approval.

  • AI Agents

    Scoped AI agents with defined tools, least-privilege credentials, human approval on write actions and a full audit trail of every call.

Start with an assessment

The fastest way to a useful answer is a short, scoped look at what you already have.