The diagnostic examines five connected capability groups. Together, they cover the quality, location, governance, accessibility, protection and organisation of the data and content that AI systems may rely on.
Data Quality & Consistency
This group examines whether your organisation's core data can be trusted. It covers whether teams work from one agreed source of key figures, whether decision-makers rely on system data without constant double-checking, and whether processes keep records current and free of duplicates that would mislead people and AI tools alike.
- Single source of truth
- Data confidence
- Duplication and currency control
This capability is about whether your teams work from one trusted source for key business figures. When revenue, customer numbers or performance metrics come from a single agreed place, meetings start from shared facts. When they come from competing spreadsheets, time is lost reconciling versions and decisions rest on whichever number won the argument. This capability measures whether decision-makers genuinely trust the data in your core systems. Where confidence is high, people act on reports without pausing to verify them elsewhere. Where it is low, shadow checking becomes routine, decisions slow down, and any AI tool built on that same data inherits the same doubt. This capability covers the processes that keep customer and operational records current and free of duplicates. Duplicate and outdated records confuse staff, waste effort and embarrass the organisation in front of customers. For AI, they are worse still, because models and assistants treat stale or repeated records as equally valid truth.
Data Architecture & Accessibility
This group looks at where your data lives and how easily it moves. It covers visibility of your data estate across systems and locations, whether core systems exchange data without manual export and rework, and whether people who need data for analysis can readily access it without long IT delays.
- Estate visibility
- Integration readiness
- Analytical accessibility
This capability asks whether you actually know where your critical business data lives across systems, platforms and locations. Many organisations discover during AI projects that important data sits in forgotten databases, personal drives or legacy applications. Without a clear map of the estate, you cannot judge quality, protect sensitive material or plan integration. This capability covers whether your core systems can share data with each other and with new tools without manual export and rework. Well-integrated systems let information flow to where it is needed automatically. Poorly integrated ones force people to download, clean and re-upload files, introducing errors and making AI adoption slow and fragile. This capability examines whether the people who need data for analysis can reach it without lengthy IT requests. When analysts and business users have governed self-service access, questions get answered in hours rather than weeks. When every extract requires a ticket, curiosity dies, analysis stalls and AI experiments never get the data they need.
Information Governance & Classification
This group assesses whether your information is properly governed. It covers whether key data sets have named owners accountable for quality, whether information is classified and labelled so people and systems handle it correctly, and whether retention rules remove outdated and redundant material instead of letting it quietly accumulate over years.
- Ownership and stewardship
- Classification and labelling
- Retention discipline
This capability is about whether your key data sets have named owners who are responsible for their quality and use. Ownership turns data care into someone's actual job rather than everyone's vague duty. Without named stewards, errors go unfixed, standards drift and nobody has the authority to decide how data should be used. This capability covers whether your information is classified and labelled so that people and systems know how it should be handled. Labels such as confidential, internal or public guide daily behaviour and drive automated protections. Without them, sensitive material is handled like routine content, and AI tools cannot tell what they should never expose. This capability assesses whether retention rules are actually applied so outdated and redundant information is removed rather than accumulating. Old material buries current answers, inflates storage costs and increases legal exposure. Disciplined retention keeps the estate lean, meaning searches, analytics and AI assistants draw on information that is still accurate and relevant.
Permissions & Access Hygiene
This group examines who can reach your data and whether that access is controlled. It covers whether access rights have been recently audited and permission sprawl remediated, whether you know the extent of oversharing in collaboration environments, and whether access is reliably removed or updated when people leave or change roles.
- Access audit currency
- Oversharing visibility
- Leaver and role-change control
This capability asks whether you have recently audited who has access to what and remediated permission sprawl. Access rights accumulate quietly over years as people join projects and never leave them. Regular audits catch this drift before it becomes dangerous, and remediation ensures that access reflects what people currently need, not history. This capability covers whether you know the extent of oversharing in your collaboration and file-sharing environments. Links shared with everyone, open sites and inherited permissions quietly expose far more than anyone intends. AI search makes this worse, because assistants surface anything a user can technically reach, including material they were never meant to see. This capability examines whether access rights are reliably removed or updated when people leave the organisation or change roles. Leavers who keep credentials and movers who keep old permissions are among the most common causes of data exposure. Reliable joiner, mover and leaver processes close these gaps quickly and consistently every time.
Unstructured Content Management
This group looks at the documents, emails and collaboration content that make up most organisational data. It covers whether you understand the scale and state of that estate, whether sensitive content is located and protected, and whether the estate is organised well enough for AI tools to surface accurate, current answers.
- Content estate visibility
- Sensitive content control
- AI readiness of content
This capability is about whether you understand the scale and state of your document, email and collaboration content. Unstructured content usually dwarfs structured data and grows without oversight. Knowing how much exists, where it sits, how old it is and who owns it is the starting point for cleaning it up for AI. This capability covers whether you know where sensitive content sits within your unstructured data and how it is protected. Contracts, personal data, salary details and strategy papers scatter across drives and sites over time. If you cannot locate and secure them, any AI tool searching your estate may expose them in seconds. This capability asks whether your document estate is organised well enough that an AI tool searching it would surface accurate, current answers. AI assistants cannot tell a final policy from a superseded draft unless the estate makes that clear. Curated, current, well-structured content is the difference between trustworthy AI answers and confident nonsense.