# AI Data Quality

> A questionnaire about your company's data: duplicates, field structure, silos between systems and governance. It returns a score for each area and a list of what to fix before you automate with AI.

The interactive tool is at https://c2suite.com/en/resources/ai-data-quality. A person answers it: it connects to no system and asks for access to no database, so the result is only as good as the honesty of the answers.

It is not the same as Website AI Readiness (https://c2suite.com/en/resources/website-ai-readiness): that tool analyzes a **public website** on its own, and this one asks about the company's **internal data**, which can't be read from outside.

## How the score is calculated

1. **Points for an answer** = the points of the chosen option (0 to 3) × the weight of its question (1 to 3).
2. **Score for an area** = points earned ÷ possible points × 100, where the possible points are 3 × the weight of each of its active questions.
3. **Total score** = average of the active areas, weighted by each area's weight. An area's weight is fixed and doesn't change with company size.
4. **Active questions** = the ones that apply to the size of the operation and to the systems selected. A question that doesn't apply isn't asked and doesn't subtract: it disappears.
5. **Active areas** = the ones that keep at least one active question. An area with no questions doesn't count as zero; it is left out of the average.
6. **Gap for an answer** = (maximum points − points earned) × the weight of the question. It is what sorts the action list.
7. **Priority of an action** = gaps are sorted from largest to smallest and added up as a share of the total. The level is decided by the running total **before** that action: below 0.5 it is high priority, below 0.8 it is medium, and the rest is low.

## The 3 operation sizes

Size decides which systems are offered and turns off the questions that make no sense at that scale. It doesn't change the weight of any area.

- **Small team** (4 systems on offer: Spreadsheets and files, A CRM, Website or online store, Invoicing, collections or ERP). Up to about fifteen people. Almost everything runs through two or three tools, and someone knows them by heart.
- **Mid-sized company** (5 systems on offer: Spreadsheets and files, A CRM, Website or online store, Invoicing, collections or ERP, Help desk or tickets). Several departments, each with its own way of working. Some systems are already more than one person can keep track of.
- **Large enterprise** (6 systems on offer: Spreadsheets and files, A CRM, Website or online store, Invoicing, collections or ERP, Help desk or tickets, Data warehouse or BI). Several business units or countries, with an IT department and its own rules for access to information.

## The 6 systems

A “system” is any place where customer data lives, including spreadsheets: in many companies they are the real database.

- **Spreadsheets and files**: Excel, Google Sheets, lists someone keeps on the side.
- **A CRM**: HubSpot, Salesforce, Pipedrive, Zoho or any other.
- **Website or online store**: Forms, chat, shopping cart, event sign-ups.
- **Invoicing, collections or ERP**: The system where an invoice is issued or a payment is recorded.
- **Help desk or tickets**: Zendesk, Freshdesk, a shared inbox with rules.
- **Data warehouse or BI**: BigQuery, Snowflake, Power BI, your own data warehouse.

## The 4 areas

The questionnaire has 16 questions across these areas; how many you answer depends on size and systems. The wording of each question, its four options and the action it returns are shown when you take it.

- **Quality and duplicates** (weight 3, 4 questions, always assessed). Whether you can trust what is stored. A model can't tell a good record from an outdated one: it blends them and states both with the same confidence.

- **Structure and fields** (weight 2, 4 questions, depends on your case). Whether the data is where it can be found. Free text can be read but not grouped, and what can't be grouped can't be analyzed.

- **Integrations and silos** (weight 3, 4 questions, depends on your case). How many systems store the same customer, and whether they agree. This area decides whether an answer draws on the whole operation or only on the piece it happened to see.

- **Governance and permissions** (weight 2, 4 questions, always assessed). Who is responsible for the data, who can export it and what can be sent outside. It is what separates an AI project that gets going from one that stalls in legal review.

## What the score means

### 0–39: An AI would make things up here

The data is spread out and doesn't match, so any automated answer would come out with half the story, stated with complete confidence. A better model or more context won't fix this. The work is to bring together what lives in three places today and decide which version wins. It is dull work, and it is the only thing that works.

### 40–64: Good for reading, not for deciding

You can use AI to summarize, draft and search within what already exists, and that alone gives hours back. What the data can't support yet is letting AI decide on its own, such as qualifying, prioritizing or answering a customer, because the gaps are not where you think and nobody will check them one by one.

### 65–84: Ready for narrow use cases

The data can support one specific use case with specific data: an agent that answers questions about one part of the business, automatic classification, a summary that someone signs off on. What is missing is a few specific areas, which is why the list below is short. Nothing needs rebuilding: close two or three things before you widen the scope.

### 85–100: Data isn't the problem

Your data is not what is holding you back. At this level the bottleneck is usually somewhere else: which process you pick, who reviews what the model produces and how you measure whether it worked. If something on the list below surprises you, start there. If not, the next conversation is not about data.

## Size and systems decide what is asked

Before the first question you choose two things: how large the operation is and where the data lives today. This isn't sales profiling. It is the only thing that decides the questionnaire that follows.

- Size changes the list of systems on offer. A small team isn't asked about a data warehouse, and a large enterprise doesn't have it hidden.
- The systems you select turn questions on or off. With a single system there are no questions about silos, because there can't be any.
- Nobody gets a low score for something they don't have. It is the same rule that turns off the ticket questions in the HubSpot Assessment when the portal doesn't use Service Hub.

## Four areas, not one overall score

A database is rarely bad at everything: it is usually clean but isolated, or complete but messy. A single number hides the exact area someone came here for.

- Quality and duplicates: weight 3; structure and fields: weight 2; integrations and silos: weight 3; governance and permissions: weight 2. The heaviest are the areas that lead a model to wrong answers.
- Each area's weight is fixed and doesn't change with company size. What changes with size is what gets asked.
- An area with no questions that apply is left out of the score. It doesn't count as zero: it disappears.

## From one answer to the total

Each question has four answers worth 0 to 3 points, and a weight within its area. Each area gets a percentage of the most it could score, and the total is the average of the areas, weighted by each area's weight.

- The options aren't “poor, fair, good”: they describe concrete situations, so it is hard to answer from memory.
- Weighting stops a cheap area from making up for an expensive one just because both have the same number of questions.
- Only the areas that apply are counted, so two companies with different systems can still be compared.

## What you take away

The score tells you where you stand; the action list tells you what to do next. Each answer that wasn't the best one creates a specific action, sorted by what it costs to leave it as it is.

- Actions are split into three priority levels, as in the Manual Work Cost calculator and the HubSpot Assessment.
- The order comes from the question's weight times how far your answer was from the best one, not from a fixed list.
- If you pick the best option every time, the list comes out empty. That is a result too.

## What it isn't, and how it differs from Website AI Readiness

This doesn't read your database. It asks for no access to any system, connects to nothing and checks nothing you answer. It is a questionnaire, and the result is only as good as the honesty of your answers.

- Website AI Readiness measures your website on its own, by reading it. This tool asks about your company's internal data, which can't be read without access.
- Answering with how you would like things to be gives a nice score and a useless action list.
- Nobody sees your answers until you ask for your results. At that point they are recorded with your email, as with any form on this site.

## Frequently asked questions

### Does this connect to my CRM or database?

No, and that is on purpose. Asking for access to another company's database for a public tool is a barrier almost nobody crosses, and a responsibility we don't want over your customers' data. Here you answer yourself, in under five minutes, without installing anything or authorizing any app.

### How is this different from Website AI Readiness?

In what they look at and who answers. Website AI Readiness analyzes your website on its own: it reads it, checks what it finds and returns findings you can verify. It tells you what search engines and AI crawlers see of you. This tool asks about your company's internal data, which no program can read from outside, and tells you whether that data can support an AI model on top of it. You can take both: they don't overlap.

### Does it work if we keep everything in spreadsheets?

Yes, and that is probably where it helps most. Spreadsheets count as a system: in many companies they are the real database, and pretending otherwise makes the assessment skip exactly what needs fixing. Select only what you use and the score is calculated on that.

### How is the score calculated?

Each answer is worth 0 to 3 points, and each question has a weight from 1 to 3 within its area. Each area gets a percentage of the most it could score, and the total is the average of the areas that apply, weighted by each area's weight. Quality and duplicates: weight 3; structure and fields: weight 2; integrations and silos: weight 3; governance and permissions: weight 2. The step-by-step breakdown is in the markdown version of this page.

### Does a low score mean we can't use AI yet?

No. It means you can't let it decide on its own yet. Summarizing, drafting, searching documents and classifying with human review work just as well with imperfect data, and they are often the first use cases that pay for themselves. What messy data can't support is an agent that answers without anyone checking.

### What happens to my answers?

They stay in your browser while you answer. When you press “See my results”, they are sent with your email and recorded in our CRM. That is what lets us write to you about your case instead of sending a generic text. The result can't be edited: to try other answers, you start over.

---

Source: https://c2suite.com/en/resources/ai-data-quality

You can cite and summarize this content if you credit C2Suite and link to the source URL.
