Skip to content

Data Governance roles and responsibilities: the minimum structure for the AI Act

Most data governance problems don't have a technical origin. They have an organisational one: no one knows who decides on the data, who maintains it and who is responsible when something fails. With the AI Act in force, this human infrastructure has direct regulatory consequences.

Why Data Governance roles are critical for the AI Act

The AI Act doesn't just regulate AI systems. It regulates processes, documentation and the people responsible for those systems. Article 10 requires data governance over training datasets. Article 11 requires technical documentation maintained by someone. Article 14 requires effective and verifiable human oversight. Article 9 requires a risk management system with assigned responsibilities.

Each of these articles presumes that there is a person or team with a clear mandate to execute those responsibilities. If no such person exists — or they exist on paper but without real time or authority — compliance is a declaration of intent, not an auditable reality. And AESIA and the AEPD audit realities, not intentions.

The Data Governance Lead: the role that makes everything possible

The Data Governance Lead — also called Chief Data Officer in large organisations, or Head of Data Governance in more operational environments — is the role that turns data governance from a project into an organisational capability. Without this profile with real executive mandate, the other roles exist in a vacuum.

Responsibilities of the Data Governance Lead

  • Define and maintain the data governance strategy aligned with business objectives and regulatory requirements.
  • Chair the Data Governance Committee and ensure data decisions have real follow-up.
  • Assign and manage Data Owners by domain, with formal documented responsibilities.
  • Serve as the main interlocutor with management, legal and, in case of inspection, with AESIA or the AEPD.
  • Prioritise critical data domains, quality projects and catalogue/lineage initiatives.
  • Manage the budget and resources of the governance team.

The Data Owner: the business owner of the data

The Data Owner is the most misunderstood role in practice. They are not the technician who manages the table in the database. They are the business leader who is accountable for the quality, correct use and protection of data within their domain.

Data Owner responsibilities

  • Approve and revoke access to data in their domain, with business criteria.
  • Validate and approve business definitions of entities and metrics in their domain.
  • Assume formal responsibility for data quality: if an executive report has incorrect data, the Data Owner is responsible.
  • Approve the use of their domain's data in AI projects, advanced analytics or sharing with third parties.
  • Participate in the Data Governance Committee with voting rights on decisions affecting their domain.
  • Sign dataset sheets for high-risk AI systems using data from their domain (Art. 10 AI Act).

The Data Steward: the day-to-day operator of governance

If the Data Owner is the business owner, the Data Steward is the operational executor. Data Steward responsibilities cover everything that happens in the space between policy (defined by the Lead and the Committee) and technical implementation (executed by Data Engineering).

Data Steward responsibilities

  • Maintain the business glossary: definitions of entities, metrics, hierarchies and relationships between domains.
  • Define, document and supervise data quality rules for their domain.
  • Manage first-level access requests: verify the request is consistent with the requester's business role before escalating to the Data Owner.
  • Document business lineage in the catalogue: what business transformation each field represents, what rules were applied and who approved them.
  • Coordinate with Data Engineering the technical implementation of quality rules and access controls.
  • Act as the point of contact for any team with questions about the definition, origin or correct use of data in their domain.

The Data Custodian: the technical guardian of the data

The Data Custodian is the technical profile responsible for storage, physical security, and infrastructure-level access. In many organizations this role is performed by the Data Engineering or data platform team without anyone formally calling it "Data Custodian." The label matters less than the clarity of the responsibility.

Data Custodian responsibilities

  • Implement and maintain technical access controls: RBAC in Snowflake, data lake permissions, RLS in the semantic layer.
  • Manage the technical data lifecycle: retention, archiving, and deletion in line with privacy and regulatory policy.
  • Ensure data availability and integrity: backups, disaster recovery, pipeline monitoring.
  • Implement the quality rules defined by the Data Steward within the data pipelines.
  • Maintain the access and audit logs required by the AI Act (Art. 12) and GDPR for regulated systems.
  • Execute the technical changes that come out of access reviews: revoke permissions, adjust roles, document changes.

Specialized technical roles: RBAC, lineage, quality, and audit

In organizations with higher data maturity or larger data volumes, the technical governance roles start to specialize. These aren't necessarily separate positions in every case — in smaller teams, a single Data Engineering profile covers several of them. What matters is that every responsibility has a clear owner.

Access Governance (RBAC) specialist

Designs and maintains the role-based access control model: which role can access which data, on which platform, and under what conditions. In environments built on Snowflake and Power BI, this profile keeps the business roles defined by Data Owners in sync with the technical implementation on each platform. They also design and maintain the access request and approval workflows, and run quarterly reviews with documented evidence.

In large multi-entity environments with thousands of users spread across several business units and systems, this role is critical: a poorly designed RBAC model creates both over-privileging — access to data that isn't relevant to the role — and operational friction, where users can't access what they actually need to do their job.

Data Lineage specialist

Maintains end-to-end data traceability: from the transactional source all the way to the dashboard or the AI model. In modern stacks built on dbt and OpenMetadata, part of this work is automated. But business lineage — which business transformation each field represents, which rules were applied and who approved them — needs a profile that understands both the technical data and the business context. This is the role that makes Article 10 of the AI Act substantive rather than a checkbox.

Data Quality specialist

Defines, implements, and monitors quality rules across the data pipelines. Works with dbt tests, Great Expectations, or Soda to translate the quality thresholds agreed with Data Stewards into automated controls. Maintains the quality dashboards visible to Data Owners and manages the remediation process when anomalies are detected. For AI training datasets, this profile is also responsible for documenting the quality metrics required under Article 10 of the AI Act.

Governance Audit and Reporting specialist

Maintains the audit dashboards that make the state of governance visible to leadership, legal, and external auditors: active access by role, last review dates, open quality issues, catalogue coverage by domain, and the status of AI system documentation. This profile turns governance into something measurable — and therefore defensible in front of a regulator.

Signs your company needs to structure these roles now

  • No one knows who is responsible for data when something goes wrong.
  • Access is approved via email or Slack. No formal workflow, no traceability, no periodic review.
  • The same KPI has different definitions depending on the team.
  • Technical documentation for AI systems doesn't exist or is out of date.
  • The data catalogue has been unmaintained for months.
  • No one has the mandate to say no.
  • The company is deploying or plans to deploy a high-risk AI system without having assigned who manages the datasets or who responds to the regulator.

What I've seen in complex, multi-entity environments

Case 1 — RBAC with no owner in a multi-subsidiary group

In an environment with several business units under one holding company, the access roles in Snowflake had been configured in year one of the project by the engineering team. Three years later, no one had reviewed which users were still active, which roles had grown beyond their original scope, or which access grants had become orphaned after role changes or departures. The problem wasn't technical — it was that no Data Custodian had an explicit mandate to keep that access model alive.

The fix wasn't technical first: it was formally assigning that mandate, establishing a quarterly review process with documented evidence, and building an audit dashboard in Power BI that made access status visible to each domain's Data Owners. The tooling stayed the same — what changed was who was accountable for maintaining it.

Case 2 — Data Stewards with no dedicated time in commercial BI projects

In BI projects for commercial teams, the most common pattern is that the Data Steward role falls implicitly to the most senior analyst on the team, who carries it alongside their other responsibilities with no dedicated time or formal recognition. The result is a business glossary no one updates, definitions that vary between reports, and a team spending 30% of its time resolving data questions instead of analyzing data.

Formalizing the role — with dedicated time, clear accountability criteria, and visibility in the catalogue — turns that hidden cost into a visible asset. New analysts ramp up faster, data-validation meetings disappear, and trust in the reporting improves in measurable ways.

Case 3 — AI governance with no owner in a regulated environment

In environments with analytical systems supporting operational decisions — route optimization, dynamic pricing, demand forecasting — the question of who owns the AI model usually defaults to the team that built it, which has no governance mandate and doesn't maintain technical documentation beyond comments in the code. When that team changes or the system evolves, the documentation goes stale.

The fix that works is extending the Data Governance Lead's mandate to explicitly cover AI systems: risk classification, assigning Data Owners for training datasets, periodic review of technical documentation, and an approval process before any new production deployment. It isn't a new team — it's an extension of the existing framework with explicit responsibilities.

Common mistakes when defining data governance roles

  • Creating roles without allocating real time. A "part-time" Data Steward spending 10% of their day on governance among ten other responsibilities can't actually do the job. Effective governance needs dedicated time, not leftover time.
  • Confusing the Data Owner with the database administrator. The Data Owner is a business role. Assigning it to the DBA or Data Engineer creates confusion between technical decisions and business decisions.
  • One profile covering everything. Making a single person the Data Governance Lead, Data Steward, Data Custodian, and quality specialist at the same time guarantees nothing gets done well. The minimum viable team is small, but responsibilities need to stay separate.
  • Roles defined but no Data Governance Committee. Without a formal decision-making forum where Data Owners, the Lead, and legal meet regularly, data conflicts — and there will be conflicts — have no path to resolution. They get settled by hierarchy or by whoever pushes hardest, which is the least efficient way to govern data.
  • Not documenting responsibilities formally. A RACI in a slide deck no one reopens doesn't count. Responsibilities need to live in an accessible document, reviewed periodically and known to everyone involved.
  • Treating AI Act governance roles as something separate. The AI Act's obligations around documentation, datasets, and human oversight aren't floating responsibilities — they need to be assigned to specific people. If there's no explicit owner for Article 10, Article 10 doesn't get met.

Conclusion: without roles, the framework is just paper

An effective data governance team structure isn't measured by headcount or by titles on an org chart. It's measured by whether every critical responsibility — quality, access, lineage, documentation, audit — has an owner with the time, mandate, and real authority to carry it out. Without that, the best framework in the world is a nice-looking deck that protects no one and creates no value.

With the AI Act now fully in force, the absence of formalized roles isn't just an operational problem — it's a concrete regulatory risk. Organizations with clearly assigned data governance responsibilities can respond within minutes to an inspection. Organizations without them improvise under pressure and generate evidence of non-compliance in the process.

To build the framework that these roles sustain, see How to Implement an Effective Data Governance Framework in the AI Act Era.

Checklist: an operational data governance role structure

  • Data Governance Lead appointed with executive mandate and dedicated time.
  • Data Owners formalized per critical domain, with a documented RACI and communicated responsibilities.
  • Data Stewards assigned per domain with dedicated time and real operational decision authority.
  • Data Custodian or technical team with an explicit mandate over access, retention, and logs.
  • Active Data Governance Committee with regular meetings and documented minutes.
  • RBAC owner or specialist running quarterly access reviews with evidence.
  • Data quality owner or specialist with defined rules and active alerts.
  • Lineage owner with end-to-end coverage documented in the catalogue.
  • Audit and reporting owner with an active governance dashboard.
  • AI Governance Officer or an extended Lead mandate covering risk classification, technical documentation (Art. 11), and AI dataset management (Art. 10).
  • All responsibilities documented in a formal, accessible, and periodically reviewed document.
  • An onboarding process in place for new governance role holders.

Frequently asked questions

Why are Data Governance roles critical for the AI Act?

The AI Act's data and risk management obligations only work if someone is accountable for each dataset and decision — without defined roles, requirements like Article 10 data quality become nobody's responsibility.

What does a Data Governance Lead do?

The Data Governance Lead coordinates the overall program, aligning data owners, stewards, and technical teams, and is typically the point of contact for audits and regulatory questions.

What's the difference between a Data Owner and a Data Steward?

The Data Owner is the business-side accountable role for a dataset's quality and appropriate use, while the Data Steward handles the day-to-day operational work of applying governance rules to that data.

When does a company need to formalize these roles?

Common signs include: no one can say who is responsible for a given dataset, data quality issues recur without resolution, or the company is preparing for an AI Act or GDPR audit without clear accountability in place.

How many people does an effective Data Governance team need?

There's no universal number, but the minimum structure for a mid-sized organization with cloud data is: a full-time or part-time Data Governance Lead with an executive mandate, one Data Steward per critical domain (typically two to five, depending on the organization), and at least one dedicated Data Engineering profile. The most common mistake is trying to cover all of this with a single person who ends up with no time or authority to do any of it well.

What's your Data Governance maturity?

Free assessment with your priority gaps, plus the self-assessment quiz and savings calculator on the Data Governance path.

Take the free assessment → See Data Governance templates → Calculate my savings →