An introduction to Statistical Disclosure Control (SDC)

Work in official statistics has changed dramatically in recent years. Modern technology has made it possible to produce and publish much more detailed tables and data. There are also increasing opportunities for researchers to create their own tables from the underlying microdata.

The ability to publish more detailed information brings great opportunities to better inform and support policy development and research. But it also brings the risk of disclosing sensitive information about both individual people and businesses. This is especially significant as there is now greater public awareness of privacy issues and the consequences of personal information being discovered.

Policy details

Metadata item Details
Publication date:19 August 2026
Owner:Statistical Disclosure Control (SDC) Expert Group
Who this is for:Producers of statistics
Type:Guidance
Contact:SDC.Queries@ons.gov.uk

What we mean by disclosure

A disclosure occurs when a person or an organisation recognises or learns something through released data that they did not already know about another person or organisation.

The Code of Practice (CoP) for Statistics requires statistical offices to take action to reduce the risk of disclosure and protect the confidentiality of individuals, businesses or similar. However, this can conflict with the need to produce detailed information required for the public good.

Back to top of page

Why we need Statistical Disclosure Control

In recent years statistical outputs have become more detailed and timely. They are produced in digital formats and often at low geographic level. These characteristics mean that there is a risk that personal information could be discovered through these outputs if:

  • the statistics are so detailed that it is possible to identify people with unusual characteristics
  • there are enough statistics produced on the same population and they can be linked to identify individual people
  • statistics are produced for slightly different populations, which means individual people can be identified in the data by subtracting the outputs of one source from another

Statistical Disclosure Control (SDC) encompasses the methods we can use to reduce the risk of disclosive information being present in statistics to an acceptable level, while enabling producers of statistics to release the most detailed information possible.

Back to top of page

Why we need specialist guidance on SDC

The risk of disclosure cannot be fully eliminated, so the objective of SDC is to apply a sufficient level of disclosure control to ensure the risk is acceptably low and the data remains useful. Finding the balance between risk and usefulness is one of the major SDC tasks and it can be difficult to find a satisfactory compromise.

Low counts in cells, cells with dominating contributors and unique or rare records in microdata are all potentially disclosive. However, other factors such as the sensitivity of the variables are also important.

Back to top of page

The importance of maintaining relationships with respondents and data suppliers

There are legal frameworks that regulate the publication of private information. Statistical offices need to maintain good relationships with respondents and data suppliers, as they are an essential source of information on which statistics are built. Without respondents, there are no statistics. Respondents must be able to trust that their private and sensitive information is safe in the hands of statistical offices.

Moreover, the collection of information from respondents is pointless if it fails to lead to meaningful statistics that are useful either for policy or research purposes. It is the biggest challenge of SDC to reach an acceptable balance between protecting confidentiality and ensuring outputs meet user needs. If the original information has been protected to the point that the released information is of little or no use to users, then there is little point in collecting the data in the first place.

Back to top of page

About this guidance

ONS, the National Statistician and Departmental Heads of Profession (HoPs) are responsible for statistical and survey standards for the production and release of all statistics under the Code of Practice. This guidance describes the approach that data providers should follow when producing standard outputs and ad-hoc requests. It is based on methodology developed by the ONS SDC Expert Group and ONS standards that set out minimum requirements to ensure confidentiality for public release.

This guidance serves as an introduction to more detailed guidance on SDC for:

In cases where microdata from a survey are linked with administrative microdata, you should follow the principles of the guidance for microdata from social surveys, in addition to any restrictions on the use of the administrative data. If you are creating tables from these integrated data, you should follow the guidance for tables from surveys and the guidance for tables from administrative data. It may be appropriate to conduct an intruder test to assess risk in both these instances.

Data providers must comply with the Code of Practice for Statistics to avoid reputational damage to Ministers and public bodies. The Information Commissioner’s Office (ICO) can apply stringent fines for confidentiality breaches. This SDC guidance should help mitigate and manage these risks.

All SDC guidance has been approved by Office for National Statistics (ONS) methodologists before publication and applies to Analysis Function (AF) disseminated data. The guidance in this series should be treated as manuals of good practice for all government data collections and outputs. Most pieces of guidance are supplemented by case studies.

This guidance can be used for Freedom of Information (FOI) requests in conjunction with the exemptions in the FOI legislation.

The Statistics and Registration Service Act (SRSA) 2007 includes data confidentiality regulations which apply specifically to ONS. This means this guidance can be used within ONS to ensure compliance with the SRSA, as well as outside the ONS for compliance with the 2018 Data Protection Act (DPA).

The scope of this guidance

Data in tables, or ‘tabular data’

This guidance applies to tables released to the public through standard publications and customer requests. It also applies to tables released through the Freedom of Information Act (2000) or the Freedom of Information (Scotland) Act (2002).

This guidance can be used to help assess which exemptions are relevant, and which should be cited when withholding confidential statistical information. It is good practice to explain why one might withhold information, but this must be done with reference to exemptions in the relevant Act.

The guidance can be applied to tables derived from:

  • social surveys of households, social institutions (such as schools, colleges and hospitals) or individual people — examples include the Labour Force Survey and the Living Costs and Food Survey, which rely on the knowing and voluntary participation of the respondent
  • subsamples of Census or other whole population data where only a fraction of the data source (assumed to be of the order of 1% to 10% of the target population or any subgroups of the whole population) is included in the sample without the knowledge of the original respondent — the target population is defined as the set of statistical units about which information is sought and estimates required
  • business surveys, including the Interdepartmental Business Register (IDBR)
  • sample surveys, such as the Annual Business Survey where a business is the selection unit or output unit

This guidance can also be applied to frequency tables, or tables of counts, derived from registration processes or administrative sources. These may have a near-complete coverage of the population or a sub-population.

Microdata

This guidance will ensure the ONS is compliant with Section 39 of the SRSA.

Specific legislation, such as the 2018 Data Protection Act (DPA) and UK General Data Protection Regulation (GDPR), is especially relevant when data are released by departments other than the ONS, although these are also applicable to ONS data.

Social survey microdata may be lodged at the UK Data Service (formally the UK Data Archive) and other trusted environments for secondary research. These are usually ‘safeguarded’ data released under light-touch restrictions similar to those previously covered by the End-User Licence (EUL).

Back to top of page

Main steps and background

Confidentiality protection is needed for tabular outputs and microdata. This is to:

  • satisfy legal rights and obligations
  • satisfy national and international standards for statistics
  • help maintain public trust and cooperation

Main steps

The production of anonymised outputs follows a set of main steps, from obtaining data to publication.

 

Step 1: Determine user requirements

It is important for a disclosure risk assessor to understand why users need the statistics and how will they be used.

Step 2: Understand characteristics of data and outputs

It is crucial to understand the underlying data to support decisions about how to apply disclosure control. This includes understanding the data source, how the outputs are created, and any quality issues with the data.

Step 3: Check for possible disclosure risks

Identify situations where disclosure is likely to occur. If disclosure is not likely, the data can be published. If disclosure is likely to occur, risk assessors should continue to the next step of the process.

Step 4: Check for legal and policy breaches

Ensure you understand:

  • which legislation is relevant to the data outputs
  • any policy that is in place — remember this may not be enshrined in legislation, but can help guide best practice

If you think any breaches are likely, disclosure control methods can be used to effectively manage the risks.

Step 5: Decide on the SDC method to use

Implement and publish information about both the chosen SDC methods and dissemination of the statistics. The use of disclosure control methods should involve the application of standard tools where appropriate. When looking at published statistics, users should be aware that:

  • the dataset has been assessed for disclosure risk
  • methods of protection may have been applied

For quality purposes, users should be provided with an indication of the nature and extent of any modification due to disclosure control methods. But the level of detail made available should not be enough to allow the user to recover true values where they have been protected. This may mean, for example, that the data provider may refer to the specific SDC method – or methods – but may decide not to share parameters if they were likely to significantly increase disclosure risk.

Step 6: Publish outputs

When SDC methods have been applied, the outputs can be published.

For survey data, different approaches may be appropriate for weighted and non-weighted data. Non-weighted outputs will display the actual value (which is potentially more disclosive), whereas weighted data will be summed up to reflect the population.

In conclusion to this section, if a microdata file is to be released as safeguarded data (equivalent to under an EUL), the data must be protected so that an intruder would not be able to identify a person, family, household or business, either directly from the data or by using other information in the public domain. The microdata need to include enough detail to meet the requirements of most users. Overall, the disclosure protection process must not impose an unreasonable burden on the producers of statistics.

Back to top of page

Intruder scenarios

People who try to identify individuals, businesses and other units – either on purpose or inadvertently – are known as intruders (or attackers). Intruders can be:

  • motivated — such as journalists looking for a story about someone
  • non-malicious — such as a friend or colleague who sees an output and realises they can identify someone in the data (this is known as spontaneous recognition)

Usually, a small number of intruder scenarios are considered when applying disclosure control.

It can be useful to consider possible scenarios where disclosure of data can occur. For tables based on surveys and administrative data, the main issues are identifying a member of the dataset followed by attribute disclosure. This is where additional information relating to an individual or business can be discovered.

Each type of output has its own set of common intruder scenarios.

Frequency tables

Respondents that are rare or unique in the data can be identified by:

  • use of external knowledge
  • someone who knows a particular respondent is in the dataset, including self-identification
  • someone with the same or similar characteristics

The risk is higher if there are cells with low counts, including zeros.

Magnitude tables

Tables of values (such as business turnover and number of employees) can be determined fairly accurately if:

  • they are based on a small number of contributing units
  • the value of one of the units dominates the cell total

Microdata

There is a lot more detail in record level data when compared to tables. Someone could use known information (such as age, sex, or occupation) about a respondent to try and discover more sensitive information contained in the record about that person.

Back to top of page

General guidance for standard outputs

When applying disclosure control, you should consider how an intruder could attempt to identify a respondent in an output that might contain potentially disclosive data.

Disclosure control techniques suitable for defined outputs are described. These need to be considered as part of the wider context of the data release and any exemptions that may be applicable.

Tabular data

Tables produced from surveys or admin data are usually for a large geographical area such as a Region or Country. For some variable combinations, outputs at Local Authority District level will be available. Demand for lower geographies is becoming greater and, with smaller counts in smaller areas, this provides challenges for protecting confidentiality.

The responses being protected may be:

  • individual people
  • families
  • households
  • any other unit whose confidentiality should be protected

Examples of protection methods

This may apply where there is one unit contributing to a cell.

You should first consider whether it would be possible to combine categories within one or more variables so that unsafe cells might be grouped together. This should always be the initial approach, to be applied alongside consideration of the user requirements. Variable categories can be combined, or variables removed until only safe cells remain.

Sometimes ungrouped categories are necessary. In these cases, an approach could be to suppress unsafe cells. This involves replacing the value within the cells with a ‘c’ or some other symbol.

This technique can be used for frequency tables when there are low counts in the table (cells of size 1 and 2 are usually risky). Where the sample size of a total or sub-total is one or two, suppress the whole row or column to which the total refers, including any zero cells. You could also combine neighbouring categories.

For magnitude tables, both of the following must apply:

  • there must be at least ‘n’ enterprise groups in a cell – this is known as the threshold rule
  • the total of the cell minus the largest ‘m’ reporting unit (or units) must be greater than or equal to p% of the value of the largest reporting unit – this is known as the p% rule

The values of the ‘p%’ and minimum threshold parameter ‘n’ and ‘m’ should remain confidential, since knowledge of these values reduces the protection. The choice of ‘p’, ‘n’ and ‘m’ would usually be decided by the Responsible Statistician.

Typical examples would be:

  • 2,3,4,5 (for ‘n’)
  • 2,3 (for ‘m’)
  • 10%, 15%, 20% (for ‘p’)

In the case of business data, there may be legal restrictions on the choice of these parameters, particularly the threshold ‘n’. The Statistics of Trade Act (1947) mandates that there must be a minimum of 5 returns contributing to any statistic.

The suppression of unsafe cells is known as ‘primary suppression’. Once these cells have been suppressed, other cells must also be suppressed to prevent the values of the unsafe cells being calculated by subtraction from the marginal totals (row and column totals, or other subtotals) of the table. These suppressions are known as ‘secondary suppressions’.

When several tables are published with common variables, disclosure by differencing one table from another can occur where the variable classifications are slightly different. An example would be where one table uses variable classifications of ‘0 to 4’ and ‘5 to 9’, and another uses classifications of ‘0 to 5’ and ‘6 to 10’. In this case, information on the 5-year-olds could be derived. If the same geographies and variable breakdowns are always used (at each level of detail) there should be no differencing problems.

Rounding is commonly applied to counts from business surveys. This is often published alongside magnitude tables, if table redesign does not solve all problems. Controlled rounding is the preferred method, as it preserves the additivity of the table.

It is possible to release percentages and rates, as long as it is not possible to work out the values of those cell counts that would otherwise need to be protected.

If outputs are rounded, any published percentages and rates should be calculated from these rounded values.

Cell suppression does not always provide enough protection in unweighted tables. If unweighted sample base numbers are essential, they should be conventionally rounded to base 10.

By raising the geography level, small cell counts at lower geographies will be grouped together to reduce the likely number of risky cells.

Tables produced from administrative data

Similar rules apply to tables produced from administrative data, or ‘admin data’. These outputs are typically frequency tables, with low counts the major concern. Low counts in the margins of a table are a particular issue because:

  • the row or column contributing to the margin will consist of small counts
  • the marginal total will highlight the small number of contributions to a single variable and give encouragement to an intruder to investigate further

Administrative data do not always have the protection given by a survey, as in theory they may show information relating to a complete, well-defined, population. Many of these outputs relate to sensitive information such as health data and the level of risk is closely related to the population size. Outputs can therefore be separated into low, medium or high risk depending on the underlying population.

Methods of protecting administrative data outputs are like those used for survey data. Table redesign is recommended, with controlled rounding and suppression as alternatives (if the number of unsafe cells is low).

The k-Anonymity approach is sometimes applied. This is a measure used within SDC to help assess whether there is sufficient uncertainty within a microdata file. There should be at least ‘k’ records with the same combination of attributes.

Microdata

Microdata from surveys can be categorised as to whether they disclose private or personal information, or whether they would do so in combination with other data sources.

ONS microdata disclosing private or personal information

ONS microdata that discloses private or personal information can be made more available to Approved Researchers if the data fulfil one of the exemptions in the SRSA Section 39 (4). These exemptions occur when the information:

  • is required or permitted by any enactment
  • is required by The UK Statistics (Amendment etc.) (EU Exit) Regulations 2019
  • is necessary to enable the Board to perform any of its functions, or to help them with this
  • has already lawfully been made available to the public
  • is made available following an order of a court
  • is made available for the purposes of a criminal investigation or criminal proceedings – this applies whether the investigation or proceedings are taking place in the United Kingdom or elsewhere
  • is made available with the consent of the person to whom it relates
  • is made available to an Approved Researcher

The Data Protection Act (DPA) regulates the use of personal information for all data with authorisation for the release of private data being decided by the Data Controller. This applies to all data, ONS or non-ONS.

Personal data must be:

  • used fairly and lawfully
  • used for limited, specifically stated purposes
  • used in a way that is adequate, relevant, and not excessive
  • accurate
  • kept for no longer than is absolutely necessary
  • handled according to people’s data protection rights
  • kept safe and secure
  • not transferred outside the UK without adequate protection

A dataset is not considered to be ‘personal information’ under the SRSA if there is negligible risk of disclosure resulting from the dataset or in combination with any other data or information in the public domain. However, a dataset may be personal data under the DPA if there is a reasonable chance that an intruder could use privately held data to disclose information about an individual case in the dataset. In those cases, an EUL can be used, which is equivalent to the use of the term ‘safeguarded data’. Note that this assessment is not related to the likelihood of an intruder existing, or being motivated. The assessment relates to the likelihood that any intruder would have a reasonable chance of finding out something about an individual case.

Microdata that may be disclosive in combination with privately held data

Examples of privately held information include the following:

  • an inquisitive individual who could try to link their own knowledge base of friends, colleagues and acquaintances with published microdata – another example could be a club official who has a list of members with accompanying information that has been collected on a membership application
  • personal knowledge of friends or neighbours, informally gathered over a period of time – it is conceivable that this could be used alongside a published dataset to disclose additional details
  • an individual who may know an acquaintance is in a dataset, which could encourage them to search the data for this person with a hope that additional information could be discovered – this is known as ‘response knowledge’

Protection against disclosure through combination with privately held data is provided by the restrictions of the EUL. One of the EUL conditions is that users agree to ‘preserve the confidentiality of, and not attempt to identify, individuals, households or organisations in the data’. This means the EUL is also considered to provide adequate protection against disclosure resulting from the use of other data, such as private databases.

Public use files

If data were to be released for public use, data providers would need to consider the possibility that an intruder might have access to a private data source or to privileged information which could be matched with the microdata to enable someone to be identified. In order to protect microdata to this level, the usefulness of the data is likely to be seriously reduced.

These datasets may be of little practical use for research and are likely to be used largely for teaching and training purposes. An example of a published dataset of this type is the teaching file released from 2011 UK Census, later repeated for 2021. Access the 2021 teaching dataset.

Back to top of page

Implementation and evaluation

Tabular data

The Code of Practice v3.0 (Section 4.5, paragraph 20) states that official statistics producers must “protect the confidentiality of individual and business information when producing statistics”. This is to be balanced with the requirement from the UN Fundamental Principles of Official Statistics that statistics “meet the test of practical utility” (Code of Practice v3.0, page 6). Ultimately, the aim is to apply SDC which results in a satisfactory balance in the output between risk and utility.

Providers should be transparent on the disclosure protection method used and the effect it may have on quality. This is outlined in section 4.5, paragraph 20 of the Code of Practice for Statistics 3.0.

Providers should also include footnotes for tables in releases. This is shown in the guidance on tables produced from survey data. No specialised software should be needed to implement this guidance for social surveys and subsamples.

The Tau Argus open source software is available to implement cell suppression and controlled rounding for business surveys. Tau Argus is the most thorough of several disclosure control packages and is the international standard for protecting confidentiality.

Exemptions

There are situations where this guidance may not apply, such as when respondents agree a waiver. If exemptions from the standard confidentiality requirements are legislated these must be recorded and reported to the Information Asset Owner.

Disclosure control is not an exact science. Methods other than any standard ones described in this guidance may be used for confidentiality protection if one of the following applies:

  • they can be shown to provide equivalent protection to the standard method
  • higher levels of protection are required because of special circumstances relating to the data

You can seek further advice from the SDC Expert Group at the ONS by emailing SDC.Queries@ons.gov.uk.

Decisions on lower levels of protection

If none of the exemptions are applicable but a data provider wishes to release more information than would be consistent with this guidance, approval is required from the Head of Profession for Statistics or the Information Asset Owner. The SDC Expert Group can also advise.

Releasing disclosive tables

Particular procedures on access and user conditions must be followed when tables are potentially disclosive. The same process should be followed as for disclosive microdata. The tables can be released to researchers under special conditions agreeable to the Information Asset Owner or the Responsible Statistician.

Microdata

The ONS procedure for releasing microdata is governed by a panel from different business areas within ONS and includes external representation. The ONS SDC Expert Group will carry out a risk assessment of any proposed release.

To help the Expert Group make this assessment, the data provider will need to complete the social survey microdata checklist. The checklist gives an opportunity to specify the main variables, together with information about how they may be protected. The risk assessment, which may include advice on further protection needed, will be attached to the release application.

The SDC Expert Group can advise the data provider on how to ensure that data are not personal information. However, it is the responsibility of the data provider to ensure that disclosure control advice has been correctly implemented when microdata are made available to other organisations under the EUL. If additional variables or additional details are required for a future release of the data, the SDC Expert Group will need to be informed.

In practical terms, there is little difference between disclosure control procedures for publishing ONS microdata and non-ONS microdata. The same risk assessments should be carried out. The aim is to publish data that are highly useful and of low risk. One difference is that the SRSA is applicable only to ONS outputs, so other departments need to establish the legislation relevant to them.

The DPA gives guidance about sensitive variables for both ONS and non-ONS outputs, but lists no penalties for releasing confidential data. Differences in the SRSA and DPA may lead to a (contested) assumption that the microdata release position is less restrictive for GSS (but non-ONS) data as there are clearly defined penalties under the SRSA. From the user perspective there will probably be little difference in the level of detail in ONS and non-ONS outputs. For all outputs, respondents may be reluctant to respond to surveys if they suspect the microdata are not assessed rigorously before being released.

Intruder testing

The ICO Anonymisation Code of Practice recommends carrying out a test to see what an intruder could find out from the data before microdata are released. This is not always necessary but should be considered, balancing the sensitivity of the data with the resources required for intruder testing.

The testing involves simulating intruder conditions, by giving several volunteers access to the data and relevant internet sources. The volunteers should ideally have some knowledge of the data. If they can correctly identify individuals, households or businesses (and previously unknown attributes) within the data in a pre-defined time, then further protection is required. The intruders are required to sign a confidentiality declaration.

Read more about intruder testing.

Mixture of pre and post tabular processes

Some SDC methods are applicable to both tables and microdata. Some Government Departments and other data suppliers allow public access to interactive tabulation tools that can be used to generate user defined tables. It is likely that this form of dissemination will become more popular in the future allowing detailed tables to be extracted. There will be a very large number of possible tables using potentially every combination of variables in the dataset. Techniques must be applied both to the:

  • microdata before the tables are generated — this is to protect unusual characteristics
  • tables themselves — this is to protect against attribute disclosures and disclosure by differencing where users can build similar tables

This would ensure that these tables are not disclosive.

Read more about the methodological challenges of protecting outputs from a flexible dissemination system.

Back to top of page

Responsibilities

Each publication should have an associated Responsible Statistician. They are responsible for:

  • confidentiality protection of released data
  • ensuring that standard disclosure control methods are applied
  • ensuring any other special circumstances are considered

The disclosure control method will usually be signed off by the Head of Profession for Statistics or their delegated representative. This may not always the same person as the Responsible Statistician. For certain releases, this may be the National Statistician or the Chief Statistician in a devolved administration.

Day to day management of disclosure control for data release may be delegated to:

  • output managers
  • data managers
  • others responsible for the confidentiality guarantee pertaining to outputs

This applies whether the data are released through the Analysis Function, or by others using data from this source.

Back to top of page

Further reading

This guidance complements standalone Government Statistical Service (GSS) advice on topics such as:

There are several relevant documents which discuss data confidentiality in general terms, including the:

Back to top of page

Help and support

The SDC Expert Group at the ONS can help and offer advice throughout the GSS to data processors and data providers, where necessary for all data releases. You can contact the Expert Group by emailing SDC.queries@ons.gov.uk.

Back to top of page