Unity Catalog Permissions: Planning Roles and Data Access
Who can read sales data, who can change tables, and which pipeline can write to production? Without a shared permissions model, these questions quickly become individual judgment calls. That makes audits harder, slows down approvals, and leaves teams with access rights whose original purpose nobody can explain.
Unity Catalog permissions let you control access to data objects centrally in Databricks. But running individual GRANT statements is not enough. You need a structure that connects business ownership, technical identities, and verifiable rules. Here is how to build it.
In short: A clear Unity Catalog permissions model uses catalogs and schemas as access boundaries, groups for people, and service principals for automated processes. Grant only the privileges each task requires, account for inheritance, and regularly check both direct and indirect routes to data access.
How do Unity Catalog permissions work?
Unity Catalog permissions determine which identities can access data objects and which actions they can perform. Accessing a table requires both the relevant data privilege and usage privileges on its parent catalog and schema.
The basic hierarchy is:
- Metastore: the top-level administrative container for Unity Catalog objects.
- Catalog: an organizational boundary, such as an environment or business domain.
- Schema: a subdivision of a catalog, for example by data product or processing stage.
- Data object: a table, view, or volume, each with applicable privileges.
Reading a table normally requires USE CATALOG, USE SCHEMA, and SELECT. The two usage privileges alone do not grant access to table contents.
Inheritance matters: SELECT granted on a catalog or schema applies to its tables and views, including future objects. A narrow table grant therefore does not restrict broader access already in place. Always check parent-level grants and all group memberships.
How should you structure catalogs and schemas?
Catalogs and schemas should reflect stable ownership and access boundaries, not every change to the organizational chart. Separate areas when their sensitivity, accountable owners, or permitted audiences differ.
Make catalogs deliberate security boundaries
For a smaller platform, catalogs named dev, test, and prod may work well. More distinct business domains may call for combinations such as prod_sales and prod_finance. The naming pattern matters less than whether the structure makes access decisions understandable.
Before implementation, ask:
- Which teams own the data?
- Which datasets must not share an access policy?
- Which workspaces should be allowed to use the catalog?
- Who approves structural and access changes?
Workspace bindings can restrict a catalog to specific workspaces. They complement object permissions rather than replacing them.
Design schemas around access patterns
Within prod, schemas such as sales_raw and sales_reporting could have different access policies. Schema-wide read access may be appropriate if a reporting schema contains only approved data. If it also contains sensitive employee information, that boundary is too broad.
Before granting privileges at a parent level, decide whether future objects should automatically become accessible to the group.
How do roles, groups, and service principals fit together?
Business roles describe responsibilities; groups collect the privileges people need to perform them. Service principals are technical identities for automated workloads and should receive permissions independently of personal user accounts.
Prefer groups over permanent individual grants
Use account groups, ideally synchronized from your central identity provider. Workspace-local groups are not a suitable substitute for this Unity Catalog group model.
A simple pattern might distinguish:
sales_analysts: read approved sales data.sales_engineers: process and maintain defined datasets.sales_data_owners: take business responsibility and approve requests.
These names are organizational conventions, not built-in Unity Catalog roles. Business accountability also does not automatically require technical ownership of every object.
Scope technical identities carefully
A production ingestion pipeline should run as a designated service principal. Separate technical identities by environment and responsibility so that development processes do not automatically receive production write access.
Document each service principal's purpose, responsible team, required objects, and authentication method. Prefer short-lived authentication over long-lived personal tokens. Permission to start a job and the data privileges of its execution identity are separate concerns: check the configured run-as identity too.
Which privileges does a team actually need?
A team needs only the privileges required for its specific tasks. Reading data, modifying records, creating objects, and managing permissions are different capabilities that should be assessed separately.
For an analyst group, access to one approved table could look like this:
GRANT USE CATALOG ON CATALOG prod
TO `sales_analysts`;
GRANT USE SCHEMA ON SCHEMA prod.sales_reporting
TO `sales_analysts`;
GRANT SELECT ON TABLE prod.sales_reporting.monthly_revenue
TO `sales_analysts`;
This example grants targeted access; it is not a complete security configuration. If the group already has SELECT on prod, its access remains broader.
For a pipeline, distinguish source and target privileges: SELECT for reading, potentially MODIFY for changing data, and the appropriate creation privileges when new objects are needed. Required catalog and schema usage privileges still apply. The exact combination depends on the operations performed.
Restrict ownership and MANAGE carefully. Although MANAGE is not a read privilege, it allows grant management and can therefore be used to grant data access. Include these administrative capabilities explicitly in security reviews.
How can you verify access reliably?
A reliable access review combines existing grants, group memberships, administrative capabilities, and actual usage. A list of direct table grants alone does not reveal every effective access path.
Check regularly:
- Inheritance: Which privileges come from the catalog or schema?
- Identities: Do group members and technical accounts still need access?
- Administration: Who owns objects or can change grants?
- Execution: Which identities actually run jobs?
- Alternative paths: Do cloud IAM permissions allow direct access to underlying storage?
Use tools such as SHOW GRANTS, relevant information schema views, and audit logs. Respect their visibility limits: a query run by a restricted identity does not necessarily provide a complete permissions inventory.
Also test both allowed and prohibited actions using representative identities. If access persists after revoking a table grant, investigate other grant paths before assuming a technical fault.
How we approach implementation
A workable model starts with concrete tasks and ends with verifiable access tests. In our Databricks consulting engagements, we translate these requirements into practical rules together with platform teams, business owners, and IT security.
From data inventory to a permissions matrix
- Establish the baseline: Inventory data objects, sensitivity, workspaces, groups, and technical identities.
- Define boundaries: Structure catalogs and schemas around environments, ownership, and access patterns.
- Build the matrix: Record object scope, privileges, approvers, and purpose for each group and service principal.
- Control changes: Deploy grants through versioned, reproducible processes where possible; justify exceptions and set expiration dates.
- Test effectiveness: Verify expected access and explicitly prohibited actions, then document the results.
Plan for ongoing operations
The handover should cover employee onboarding, role changes, and departures, as well as retiring technical identities. Define review triggers and a risk-based review schedule. Every exception needs an accountable owner and either an expiration date or a mandatory review date.
The outcome is more than a permissions spreadsheet: it is a process that continues to work as teams change and new data products appear.
Your next step
Start with one clearly scoped data product and use it to validate your structure, groups, and technical identities. Extend the pattern to other areas only after that validation.
Need help designing your model or reviewing existing access? Ailio's teams in Bielefeld and Hamburg support you with Databricks consulting, from the permissions matrix to a verifiable implementation.
