Data

The pre-project data quality checklist

Dashboards, migrations and AI projects don't usually fail at the clever end. They fail because nobody checked the data first. Run these ten checks before you spend a pound on anything downstream.

Why this checklist exists

Every data disaster we've been called to rescue shares an origin story: the project started at the shiny end — visuals, features, models — and discovered the foundations mid-build, at change-request prices. An afternoon with this checklist is the cheapest insurance in the industry. Score each item red, amber or green; anything red gets fixed before the project starts, not during it.

The ten checks

  1. Ownership. For each source system: who owns the data, by name? Data nobody owns is data nobody will fix. No owner, no project.
  2. Access. Can you actually get the data out — an API, a database connection, a scheduled export? "Sarah downloads it monthly" is not access; it's a bottleneck with a holiday allowance. (Where no route exists, extraction engineering is the fix — but know that before you budget.)
  3. Uniqueness. Pick your core entities — customers, products, jobs — and count duplicates. "Acme Ltd" vs "ACME LIMITED" vs "Acme (UK)" is the classic; anything above a low single-digit percentage will poison every join and every AI answer downstream.
  4. Completeness. For the fields your project depends on, what percentage is actually populated? A margin dashboard over 60%-complete cost data is a fiction generator with a colour scheme.
  5. Consistency. Do definitions match across systems — is "revenue" the same calculation in the accounts package and the sales tracker? If two systems disagree today, your dashboard will simply publish the argument.
  6. Freshness. How stale is each source, and how stale can decisions tolerate? Real-time is rarely needed; known latency always is.
  7. History. Does the source keep the past, or overwrite it? Trend reporting needs history; if systems overwrite, you need to start capturing snapshots now, because history can't be backfilled.
  8. Format sanity. Dates as dates (and one convention), numbers as numbers, categories from a controlled list rather than free text. Free-text category fields are where analytics goes to die.
  9. Volume reality. Row counts and growth rates, honestly measured — they drive the backend decision and the licensing arithmetic before anyone commits to either.
  10. Sensitivity. Which fields are personal data under UK GDPR, who's allowed to see them, and does the project design respect that from day one? Retrofitted privacy is expensive; breached privacy is worse.

Where we fit in

Power Analytix runs exactly this assessment as a fixed-price discovery — every red item costed, every fix sequenced — scoped in writing, priced fixed, delivered by senior engineers. If you'd rather have it done than read about it, book a free scoping call.

What to do with the score

Mostly green: proceed, and enjoy being the rare project that finishes on budget. A few ambers: fix them inside the project scope, knowingly and priced. Any reds on ownership, access or uniqueness: stop — fixing those is phase one, whatever the original plan said. Teams that skip this step don't avoid the cost; they pay it later with interest, mid-project, when it's least affordable. The checklist is free. Use it ruthlessly.

Want this handled rather than explained?

Free 30-minute scoping call with a senior consultant. Bring the problem — leave with an approach and a price range.

Book a free scoping call +44 (0)20 0000 0000