Many AI projects start with the wrong question. The question we ask is usually, “Which model should we use?” but in my experience, the more important question is “Can we trust the data we are about to give it?” . Before a model is trained, deployed or evaluated, a series of decisions has already shaped whether that AI project will succeed or fail. Those decisions often sit far away from the model itself. They live in source systems, reporting tables, data pipelines, business definitions, manual workarounds, governance processes, and the semantic layer where an organisation decides what its data actually means.

This is why many AI projects do not fail at the modelling stage.

They fail much earlier.

They fail when key metrics are inconsistently defined.

They fail when data lineage is unclear.

They fail when manual processes are invisible.

They fail when ownership is missing.

They fail when business teams and technical teams are solving different problems.

AI does not start with the model

AI depends on foundations.

It depends on reliable data, clear definitions, consistent business rules, documented lineage, ownership, governance, and context. The model may be the visible part of the project, but it is not the only part that determines success.

In practice, AI is often treated as a modelling problem. The conversation can quickly move towards tools, algorithms, agents, prompts, automation, and deployment. But in many organisations, especially those operating in regulated or service-driven environments, the hardest work happens much earlier.

It happens in source systems.

It happens in reporting tables.

It happens in business definitions.

It happens in data pipelines.

It happens in old reports where logic has been embedded for years.

It happens in conversations with teams who understand why certain data is captured in one way and not another.

It happens in the semantic layer, where an organisation decides what its data actually means.

This part of the work is not always visible, but it is critical.

A model does not automatically understand the organisation it is being used in. An AI agent does not magically know which fields are reliable, which definitions have changed, which records are incomplete, or which processes sit outside the system. A large language model can produce fluent responses, but fluency is not the same as truth.

The model can only work with the data, context, and instructions it is given.

If the data foundation is weak, the AI output will reflect that weakness.

If the definitions are unclear, the AI output may carry that confusion forward.

If the organisation has not agreed what the data means, the AI system may produce answers that sound confident but are built on unstable assumptions.

That is why AI readiness starts before model training. It starts with understanding the data.

The hidden problem of “good enough” data

One of the biggest risks in AI adoption is assuming that data used for reporting is automatically ready for AI.

Many organisations already have dashboards, reports, KPIs, extracts, spreadsheets, and performance packs. Because these outputs exist, it can be tempting to assume the data is mature enough for AI.

But data that is “good enough” for a dashboard may not be good enough for AI.

A dashboard can sometimes survive a messy definition. An AI system may amplify it.

For example, a KPI may have been reported for years, but the logic behind it may include exclusions, manual adjustments, historical assumptions, or business rules that only a few people understand. A dashboard user may be familiar with those limitations because the report has been used in a particular business context for a long time.

An AI system does not have that informal knowledge unless the context is made available. This is where problems begin.

A field may be populated, but does anyone know whether it is still used correctly?

A status may exist, but does every team apply that status in the same way?

A customer, property, asset, case, or transaction may be marked as active, but what does “active” actually mean?

A metric may appear in several reports, but is it calculated the same way everywhere?

A process may look automated from the data, but are there manual steps happening outside the system?

These questions matter because AI systems can make poor data look more authoritative than it really is.

If the data contains inconsistent classifications, incomplete records, duplicate logic, or unclear business rules, the AI system may not challenge those issues. It may simply use them.

That is where the risk sits. The issue is not just whether the data exists. The issue is whether the data is understood.

Governance is not bureaucracy. It is risk control

Data governance is often misunderstood as documentation, process, and delay. In fast-moving AI conversations, it can be seen as something that slows innovation down, but in practice, governance is what makes innovation safer.

Good governance answers the questions that every AI project eventually has to face

  • Where did this data come from?

  • Who owns it?

  • How has it been transformed?

  • What does this field mean?

  • Which business rules have been applied?

  • Are there known data quality issues?

  • Who is responsible if the output is wrong?

  • Can the decision be explained?

These questions are not administrative extras. They are fundamental to trust.

An AI system built without governance may still produce outputs. It may even produce impressive outputs. But if the organisation cannot explain how those outputs were created, whether the underlying data is reliable, or what assumptions sit behind the result, then the AI system becomes difficult to trust.

This matters because trust is not created at the user interface. It is built through the full data journey.

It is built when definitions are clear.

It is built when lineage is visible.

It is built when ownership is assigned.

It is built when quality issues are known rather than hidden.

It is built when business and technical teams agree on what the data represents.

Without that foundation, AI can quickly become another layer of confusion on top of already inconsistent data.

The semantic layer is where meaning becomes operational

This is where the semantic layer becomes critical. A semantic layer is not just a technical convenience. It is where business meaning becomes reusable, governable, and operational.

It connects raw data to the language of the organisation. It defines metrics, relationships, calculations, hierarchies, and business rules in a way that can be consistently reused across reports, dashboards, analysis, and potentially AI use cases.

This matters because AI does not only need data. AI needs context.

If every team has a different version of “active customer”, “completed case”, “available property”, “compliant asset”, “high priority request”, or “current record”, then AI is not starting from a stable foundation. It is starting from conflicting interpretations.

The semantic layer helps reduce that ambiguity. It gives organisations a place to agree what important concepts mean. It helps ensure that when a metric appears in a dashboard, a report, or an AI-powered experience, the logic behind that metric is consistent. It also helps technical teams avoid rebuilding definitions repeatedly across different tools and projects.

For AI adoption, this is powerful.

A well-designed semantic layer can help organisations move from scattered data outputs to shared understanding. It can give AI systems better governed context to work with. It can also give business users more confidence that outputs are based on agreed definitions rather than isolated assumptions.

This is why I believe the semantic layer will become even more important in the AI era. Not less. As AI becomes easier to access, the quality of organisational meaning becomes a differentiator.

AI agents are only as good as the data behind them

The rise of AI agents makes this even more important. AI agents are often presented as tools that can take action, answer questions, retrieve information, automate workflows, or support decision-making. But an agent does not magically understand an organisation. It depends on the systems, data, instructions, permissions, and context it has access to.

If the backend data is poor, the agent’s output will reflect that.

If the definitions are inconsistent, the agent may give inconsistent answers.

If the source systems are incomplete, the agent may miss important context.

If the data is not governed, the agent may surface information that is outdated, inaccurate, or inappropriate.

This is why the phrase “your AI agent is only as good as the data behind it” is not just a catchy line. It is a practical warning.

The same applies to LLMs.

Large language models can be powerful, but they do not remove the need for data engineering, governance, architecture, or business alignment. When organisations connect LLMs to internal data, documents, knowledge bases, reports, or operational systems, the quality of those underlying sources becomes even more important.

AI may change the interface.

It does not remove the foundation.

Stakeholder alignment matters before technical delivery

Another reason AI projects fail early is that organisations sometimes move into technical delivery before agreeing on the business problem.

A model cannot fix an unclear objective.

If the business question is vague, the model output may only make the confusion look more sophisticated.

Before any AI project begins, business and technical teams need to agree on what problem they are solving, what decision the AI will support, what success looks like, and what risks need to be managed.

For example, is the goal to reduce manual effort?

Improve forecasting?

Support earlier intervention?

Summarise complex information?

Identify anomalies?

Recommend next actions?

Improve customer experience?

Each of these goals requires different data, different governance, different evaluation criteria, and different levels of human oversight.

Without alignment, teams can easily talk past each other. Technical teams may focus on model performance while business teams care about usability, compliance, explainability, or operational confidence. Leaders may expect transformation while data teams are still trying to resolve basic definition issues.

This disconnect can create disappointment, not because AI has no value, but because the organisation skipped the alignment work required to make AI useful.

Before model training begins, ask questions

Before rushing into model selection, organisations should pause and ask some practical questions.

What decision will this AI system support?

Which data sources are involved?

Who owns each data source?

Are key definitions agreed across teams?

Is the data lineage documented?

Are known data quality issues visible?

Are there manual processes that affect the data?

Has the data changed meaning over time?

Can the output be explained to non-technical stakeholders?

What could go wrong if the AI output is incorrect?

How will success be measured?

Who is accountable for maintaining trust after deployment?

These questions may not sound as exciting as discussing models, prompts, agents, or automation. But they are the questions that determine whether an AI project is built on clarity or confusion.

The organisations that answer these questions well will be in a much stronger position to use AI responsibly and effectively.

AI readiness is data readiness

AI readiness is not measured by how quickly an organisation can access a model.

It is measured by how well the organisation understands its own data.

The organisations that will get the most value from AI are not necessarily the ones that adopt the newest tools first. They are the ones that invest in strong data foundations, clear definitions, trusted pipelines, governed semantic models, and meaningful collaboration between business and technical teams.

That work can be slow.

It can be detailed.

It can involve difficult conversations about ownership, quality, legacy systems, and inconsistent reporting logic.

But it is also the work that makes AI useful.

Because before an AI system can support better decisions, the organisation has to understand the decisions it wants to support. Before a model can learn from data, the organisation has to know whether that data can be trusted. Before an AI agent can act with confidence, the systems behind it need to be reliable, governed, and meaningful.

AI may be the visible layer.

But the foundation is still data.

And if that foundation is weak, the project may fail long before model training begins.