Statistical Disclosure Control for tables produced from administrative data

This guidance follows on from the general introduction to statistical disclosure control guidance. It considers the protection of tables produced from administrative data in detail.

Related guidance is also available for:

Policy details

Metadata item Details
Publication date:19 August 2026
Owner:Statistical Disclosure Control (SDC) Expert Group
Who this is for:Producers of statistics
Type:Guidance
Contact:SDC.Queries@ons.gov.uk

The scope of this guidance

This guidance outlines the main steps to be taken in considering disclosure control, while allowing different solutions to be developed for different datasets, taking into account detailed risk assessment and the latest disclosure control methods.

Back to top of page

Risk assessment for data in tables

One major distinction between the risk assessment of tables created from survey data and tables created from administrative data is the consideration of ‘response knowledge’. This refers to the consideration that someone looking at the data may know a particular person, household or business is in the dataset.

In surveys, the sample size is usually a small percentage of the population, and someone looking at the data may have private knowledge that a particular person, household or business has responded. The cell counts are more likely to be small, which introduces a high risk that a person, household or business could be identified in the data.

Administrative data usually include most of the eligible population, and one could assume that an eligible individual is in the dataset. Though the cell counts are usually high, any small counts can lead to a high risk that a person, household or business could be identified.

Administrative data, or ‘admin data’, are data which have been collected during the operations of an organisation.

Government produces a very large amount of administrative data, providing a valuable resource if it can be used correctly. There are legal gateways which can allow accredited and approved researchers to access administrative data for research and statistical purposes. There are certain criteria to meet to ensure this can happen, including the assurance that a person’s identity cannot be identified in the information provided for research and statistics.

Administrative data are generally not collected for the sole purpose of producing statistics.

This guidance provides an overview of disclosure control techniques for all administrative data. Administrative data covers a wide area and can be collected from:

  • people or households — such as crime data, or benefit data
  • businesses — such as VAT transactions

This guidance concentrates on data collections from individuals and households.

Specific disclosure advice is available for outputs based on births and deaths registrations.

Back to top of page

Legislation and governance for outputs published from administrative data

As with survey data, there is legislation that applies to data published from administrative data sources. The legal underpinning to the release of these data is the UK Data Protection Act (DPA) 2018, which is the implementation of the UK General Data Protection Regulation (GDPR).

This Act controls how personal data are used by organisations, businesses, and the government. The government is legally bound to ensure the information collected is used fairly, lawfully and transparently.

The 2018 Act includes the following topics of personal data defined as sensitive:

  • racial or ethnic origin
  • political opinions
  • religious or other similar beliefs
  • Trade Union membership
  • health
  • sexual life
  • biometric and genetic data
  • commission or alleged commission of an offence
  • proceedings for any offence committed, or alleged to have been committed

As can be seen, sensitive personal data as defined by the DPA are wide ranging. Although information on income, wealth and financial affairs are not mentioned specifically in the Act, there may be some situations in which they could also be considered sensitive. There is little detail in official guidance as to how personal data is classified, but the variables previously described in this section are useful attributes to consider when producing tables.

The Freedom of Information (FOI) Act also can be invoked when specific pieces of information (such as a single count) are requested. These can be investigated more thoroughly than standard publications with large numbers of tables, often through use of intruder testing.

For the Office for National Statistics (ONS) and some other government bodies, current policy is for responses and released outputs from all FOI requests to be made public. This means the approaches to be followed when ensuring the confidentiality of outputs should be broadly the same as for general requests. There are several exemptions whereby FOI requests can be refused. For ONS, one of these is where the data are defined as personal information under the Statistics and Registration Service Act (SRSA) 2007 (section 39 of the SRSAsection 40 of the SRSA, and section 44 of the FOI Act). For non-ONS outputs, there may be other legislation that applies (section 44 of the FOI Act). There is also an FOI exemption that relates to personal data (section 40 of the FOI Act).

The Code of Practice for Statistics provides producers of official statistics with the detailed practices they must commit to when producing and releasing official statistics. It states that there is a legal obligation under the SRSA to ensure compliance with the Code of Practice in relation to Accredited Official Statistics (section 10 of the SRSA and section 11 of the SRSA).

Common law allows people to bring legal proceedings if the collection, use and disclosure of personal information breaches the obligation of confidence. Under the common law, anyone who receives confidential information must not disclose it without consent or justification. The common law duty of confidentiality extends to confidential information about people who are deceased.

Ultimately the Information Commissioner is responsible for ensuring good practice for publishing data is followed. They are likely to have a view on any refusals to publish on confidentiality grounds. They will need to consider the practical day-to-day aspects of disclosure control alongside the legal considerations.

The variables defined in the DPA as sensitive would likely be of high impact to a person if disclosed. In general, disclosure risk is a constant problem when personal data are released. Risk can be defined as the likelihood of a person, business, household or other statistical unit (or related attributes) being identified in a published output. The impact of any disclosure is increased if the risk involves particularly sensitive variables that could cause excessive and unnecessary harm or distress to particular individuals or groups. Disclosure of sensitive variables is also likely to be high profile and lead to considerable damage to the reputation of the data producers.

It is important to note that the term ‘impact’ refers to ‘statistical impact’. This is the effect disclosure of information will have on the person, business or household concerned. This is different to ‘reputational impact’, which occurs when a dataset is lost in transit or data are released inadvertently.

Levels of risk can be defined for outputs from administrative data by considering the geographical level of the table (or the population at risk) and the sensitivity of the variables within the table. Whether low cell counts are considered to have a potential for disclosure will depend on the level of risk associated with the table.

The sensitivity of variables is a subjective area, and the relationship between disclosure risk, variable sensitivity and the population at risk can be hard to determine. The main thing to think about is how an intruder would approach a table containing low frequencies. The data supplier has the greatest knowledge of the data, which means they are in a good position to assess how an intruder would operate. If an intruder considers the variables to be of great personal or political interest, they may invest more time in attempting to find an individual or an attribute relating to a person in the table.

Low counts in the margins of a table are a particular issue. This is not just because the row or column contributing to the margin will consist of small counts, but also because the marginal total will highlight the small number of contributions to one defined variable subgroup. This could encourage an intruder to investigate the data further.

For frequency tables, cells with low counts are likely to be the major problem. An uneven spread of frequencies throughout the table could also be concerning. In general, cell counts of 0 and 1 and rows or columns where the frequencies are concentrated into a small number of cells are the most problematic. Case Study 2 contains relevant discussion on issues around low counts.

Low risk

For some administrative based statistics, the likelihood of an identification attempt may be considered low if tables are disseminated at a high level of aggregation (such as national or regional level), with only limited tables produced from the one database. In these cases, there is usually a very small risk associated with linking between current and future releases. Statistics in this category will not usually need any protection beyond appropriate selection of variable categories and combining them where necessary. However, care should be taken where rows or columns are dominated by zeros, especially where a marginal total is 1 or 2.

Medium or high risk

Cells of size 1 or 2 may be considered unsafe for many administrative based statistics disseminated at a lower level of aggregation (for example, at small geographies or for small populations), or where many linked tables are produced from the same dataset. Care should also be taken where a row or column is dominated by zeros.

There is a similar recommendation for high-risk outputs where identification attempts may occur frequently. The impact of any successful identification would be significant for these kinds of outputs, an example being statistics on abortions. To ensure these outputs are protected, all cells of size 1 to 2 are considered unsafe and care should be taken where a row or column is dominated by zeros. High risk tables should also be closely examined if they are especially sparse, for example, if they have a low cell average and contain many zeros. Higher levels of protection may be needed for small geographical levels or for certain variables with an extremely high level of interest and impact.

Zeros

The risk levels previously discussed refer to the importance of zeros. The risk of disclosure will be greater if zeros are distributed in particular ways. If all cells in a row or column are zero apart from one, an intruder would know that all members of the row belong to a single category for the column variable. This is a form of group disclosure.

The distinction between structural and non-structural zeros is also important. Structural zeros are those where the counts cannot be anything other than zero. An example would be an impossible set of characteristics, like a mother who is 8 years old. Non-structural zeros occur because nobody with that combination of characteristics happens to be present in the population, though it would be perfectly possible for someone with those characteristics to exist. An example of a non-structural zero could be if there were no mothers aged 15 in a dataset. It is feasible for a 15-year-old mother to exist, but there just happen to be none in the data in this example. Risk in tables with zeros is generally determined by cells which contain non-structural zeros.

Population at risk

Population at risk refers to the underlying number of people that could be present in a particular cell because they share characteristics with other people in a table. The likelihood of disclosure will increase as the population at risk decreases. This means that for a smaller population at risk there is greater possibility of disclosure. For example, if a cell represented the number of women aged 18 to 34 who had given birth in a Local Authority, the population at risk would be all women aged 18 to 34 in that Local Authority.

A table is ‘unsafe’ if it contains one or more cells with an unacceptable risk of disclosure. Disclosure control methods should be used to reduce the risk by modifying these cells. Careful judgement will be required when applying any method to ensure that the most important messages are still valid.

Uncertainty

The aim for most outputs is to introduce enough uncertainty into the data to give an element of doubt to any possible disclosure. This is an alternative to the case of zero risk, where the data are protected to the extent that disclosure is not just highly unlikely but impossible. This may be a requirement for some outputs, but the changes necessary to achieve this can affect the usefulness of the data considerably.

Uncertainty relates to the extent to which cell values and apparent attribute disclosures in tables may not represent real respondents or real attribute disclosures. Published cell counts may not represent the true counts for a variety of reasons. Some of these reasons cannot be easily quantified, such as respondent error or data capture error, or mis-classification. Some can be more easily quantified, such as non-response and ‘edit and imputation’, or pre-tabulation or post-tabulation adjustment as part of disclosure control.

Factors to consider in setting the level of uncertainty include the amount of sensitivity in the data and the likelihood and impact of a real disclosure claim.

Back to top of page

Assessing the disclosure risk

There is no single, uniform approach to disclosure protection. Context is very important when you are making decisions on whether outputs pose a disclosure risk.

Determining the users’ requirements

The government produces a wide range of statistics based on administrative sources. Examples include tables of statistics concerning crime, education, benefits, and health.

Producers of statistics should design publications according to the needs of users, as a first priority. It is vital to identify the main users of the statistics and understand why they need the figures and how they will use them in detail. Read more about understanding user needs.

This is necessary to ensure that the:

  • design of the output is relevant
  • amount of disclosure protection has as little effect on the usefulness of the statistics as possible, while still being sufficient to protect confidentiality

Considering the characteristics of the data and proposed tables

The source of the data may affect the need to protect confidentiality. Sensitive variables may need special attention.

The age of the data may reduce the risk of disclosure. This is because the population of the statistics will change over time and may become less identifiable.

The quality of data may determine the extent of the need for disclosure protection. Data quality can vary considerably for many reasons, such as:

  • poor population coverage
  • poor information recording during the data processing stage
  • a need for a large proportion of imputed data
  • respondents not giving sensitive information accurately

Poor data quality may introduce uncertainty for an intruder.

Typically, statistical units are defined as people, households, or businesses. It is important to assess which units are represented in the data and need to be protected. Disclosure risks may also increase if groups of statistical units (for example, people from the same household) are represented in a table. This is because these people could be able to identify each other.

The disclosure risk for event-based data will be different than residence-based data. For example, to identify a person in a table for patients visiting a health clinic, you would need to know that the person is included in the population base for the table. In this case, this would mean knowing that a particular person has attended the clinic. The risk reduces if the population base or coverage of the table is not easily identifiable.

It is also important to consider the characteristics of the tables. Where tables are very simple and presented at a high level of aggregation (including geography), disclosure issues are less likely. However, even at a high level of aggregation, small cell counts can be a risk, and the underlying population or sub population should be considered. When tables become more detailed and the counts in individual cells are small, the risk of identification may increase and protection may be needed. If the spread of values is skewed across a table, the risk in particular cells may increase above an acceptable level.

Issues may arise with linked tables where the risk of disclosure can increase by differencing or through combining with other data. Differencing occurs when two or more tables can be considered together to deduce the value of a potentially disclosive count. For example, this may occur:

  • when tables are produced from the same dataset for two non-coterminous geographies (that is, areas with overlapping boundaries leading to small areas or slivers) — for example, Wards and Super Output Areas (SOAs)
  • where classifications differ only slightly — for example, a table that includes children aged 0 to 4 and another that includes children aged 0 to 5

Likely disclosure risks

A risk assessment should be undertaken to develop suitable confidentiality protection. This assessment should include factors such as the:

  • nature of the variables — for example, considering the risk and impact if a disclosure were to occur
  • structure of the table, meaning the number of observations in the table and their distribution

Disclosure can be quantified in terms of both risk and impact.

Disclosure risk is high when a table is designed so there are cells in the table with low frequencies, or when there are rows or columns where all the counts are in a small number of cells. Tables such as these could lead to identification of a person and maybe disclose further information about them.

The impact of any disclosure is higher if the data are sensitive and where great distress may be caused by releasing the data. For example, if a table of high risk was released giving details of treatment for mental health issues in a small geographical area, it would have greater impact than releasing a table of equal risk detailing use of local shops, as this information is much less sensitive.

Decisions on risk and impact should be made by someone with a detailed understanding of the statistics and the user need. To be explicit about the disclosure risks to be managed, you should consider a range of potentially disclosive scenarios and take action to prevent them. The risk assessment should be reviewed on a regular basis as the tolerable level of risk may change over time. The scenarios should be used to identify the specific cells that could lead to disclosure, known as ‘unsafe’ cells. Appropriate confidentiality rules should be applied to these cells. It is not possible to protect against all risks, so it is important to remember this is a risk management not a risk elimination exercise.

General attribute disclosure

Someone with knowledge relating to a statistical unit could, with the help of data from the table, discover details that were previously not known to them. Attribute disclosure includes inferential disclosure, where information about a statistical unit can be deduced, or inferred, with a high degree of confidence.

The possibility of attribute disclosure can be a concern if:

  • there is a count of 1 in a marginal total – this could be a row or column total
  • the table is distributed in such a way that one or more rows or columns are dominated by zeros

This allows an intruder to work out attribute X if they already know someone in the data with attribute Y. A zero in population data allows us to say that no-one in the population has that combination of attributes, which may narrow down the possibilities for intruders. Attribute disclosure can also occur if there is a count of 2 in a marginal total. One of the people included in the marginal total could identify the other person and potentially discover something new about them.

Disclosure risks may increase where groups of units who appear in the same table know enough about each other to identify each other and potentially discover something new. This can occur where units share characteristics or are grouped in some way. Examples could be people from the same household, or a table population that consists of a group of people who are likely to be aware of each other.

Overall, general attribute disclosure can occur in tables which are not well populated and do not have evenly distributed counts. The data owner must determine whether the sensitivity of the attribute requires the application of disclosure control.

This case study describes a situation of general attribute disclosure. This can arise when someone who has some information about a statistical unit could, with the help of data from the table, discover details that were previously not known to them.

Attribute disclosure encompasses inferential disclosure, where information about a statistical unit can be deduced, or inferred, with a high degree of confidence.

Table A1 shows female benefit claimants at a low geography along with age characteristics of the individual. A person could be identified by spontaneous recognition or through malicious intent. Attribute disclosure has occurred if someone recognises a person they may know and discovers from the table that they are claiming this benefit.

Disclosure may arise if there is a count of 1 in a marginal total (row or column). This is shown in table 1, where claimants for Sickness Benefit and Child and Housing Benefit are broken down by age bands. Anyone who knows that a person between 16 and 24 years receives a benefit would learn that it was Sickness Benefit. Attribute disclosure could occur from a count of 2 in a marginal total, where one person – or unit – may identify the other and thereby disclose further information.

Table A1: Benefit claims, by type and age (Females only) 
16 to 24 25 to 49 50 to 59 Over 59 Total
Sickness benefit 1 0 7 1 9
Child benefit 0 0 18 19 37
Housing benefit 0 12 5 2 19
Total 1 12 30 22 65

Attribute disclosure can also occur from cells with larger values, where they appear in a row or column dominated by zeros. A zero in population data allows us to say that no-one in the population has that attribute. This can be seen in Table A1, which reveals that no 25 to 49-year-old females are claiming sickness or child benefit. The risk from many zeros within tables may not be significant, but there are cases where they may need to be protected. This can depend on the distribution of the zeros and whether they dominate a row or column.

Disclosure risks may increase where groups of units who appear in the same table know enough about each other to identify each other and potentially discover something new. This can occur where either:

  • units share characteristics or are grouped in some way — an example could be people from the same household
  • the table population is some group of people likely to be aware of its other members

To protect against general attribute disclosure, at a minimum, care should be taken where rows or columns are dominated by zeros and in particular where a marginal total is a 1 or 2.

A possible solution is collapsing the age categories as in Table A1a.

Table A1a: Benefit claimants, by type and age (Females only) 
16 to 59 Over 59 Total
Sickness benefit 8 1 9
Child benefit 18 19 37
Housing benefit 17 2 19
Total 43 22 65

Much detail has been lost in Table A1a, but it still contains low counts of 1 and 2. Recoding age into 2 groups removes the attribute disclosures in the lower age groups, but low counts remain in the over 59-year-old age group. This means that combining categories has only been partially useful in protecting the table and other approaches ought to be considered.

It should be noted that in many cases combining categories can be a successful approach solving all disclosure issues, although the resulting loss in data utility needs to be balanced with the gain in data confidentiality.

Another solution is suppression as shown in Table A1b. Cells of value 0 and 1 are suppressed while other cells are selected for secondary suppression.

Table A1b: Benefit claimants, by type and age (Females only) 
16 to 24 25 to 49 50 to 59 Over 59 Total
Sickness benefit c c 7 c 9
Child benefit c c 18 c 37
Housing benefit c c 5 c 19
Total 1 12 30 22 65

Due to the nature of this table, many cells require suppression – either primary or secondary. The table is safe, but little useful information remains. The table would be more useful if the marginal totals for the two lower age groups were published, although this would require further thought as the risk of disclosure would be increased.

In the example in Table A1b, all cell counts of zero have been suppressed. However, some consideration is needed on the issue of whether always to suppress counts of zero. Zeros clearly need to be suppressed where there are enough of them to expose other counts and result in potential disclosures. It would be important to consider whether an attribute disclosure could be created by including any zeros in a table. The impact of any such ‘disclosure’ must also be assessed. Questions to consider include:

  1. How likely would it be to allow someone to find out something about an identified person with any certainty, and would the information discovered be sensitive?
  2. Would information loss be increased by suppressing zeros?

It would also be very important to consider the relationship between the usefulness of the table and confidentiality.

A further solution here is to release the data at a higher level of geography. The table could possibly be released without any disclosure control being applied, although the usefulness of the data would be compromised by producing the table at this higher level.

It may also be an option to change the age group categories, although in many cases this would not be practical as standard categories would be expected.

The motivated intruder

Data in a table can be combined with information from other sources to identify a statistical unit and disclose further details. The level of disclosure risk that is tolerable may depend heavily on the sensitivity of the data. This situation may occur when small values are reported for some cells.

Although other sources can reveal the identity of the individual, it may be the statistics that motivate the intruder to start looking and attempting to reveal what is disclosive. Official statistics should not reveal the identity of any respondents. The risk of disclosure should consider other relevant sources of information which can be combined with the statistics. The analyst does not need to consider all data sources, but should consider data which is likely to be available to third parties.

As the base population is decreased by moving to smaller geographies or sub-populations, it becomes easier to find units and discover information. It will also become more likely that an intruder will have greater confidence in any claim they might make. In a large population (for example, a country or region), the effort and expertise required to discover more details about the statistical unit may be deemed to be disproportionate for the intruder.

To protect against a motivated intruder you should, at a minimum, consider all cell counts of 1 or 2 to be potentially disclosive for geographies:

  • below Local Authority District (LAD) level – these typically range from 41,049 to 1,144,919 in England (excluding the Isles of Scilly and City of London), based on Census 2021 data
  • below Integrated Care Board (ICB) level – these range from 52,131 to 3,146,943, based on the registered population in 2022 to 2023

You should also consider all cell counts of 1 or 2 to be potentially disclosive for geographies equivalent to the size of LADs and ICBs in England.

The threshold value of 3 is chosen here to ensure that there are no instances in a well-defined geography where there is either:

  • an isolated individual
  • a member of a cell of frequency 2 that may attempt to identify the other person in the cell

Larger cell values are likely to discourage an intruder, as well as increasing the difficulty of disclosure.

As a general guideline disclosure risk increases with smaller geographies.

The situation described in this case study may occur when small values are reported for some cells. In a large population (for example, a country or region), the effort and expertise required to discover more details about the statistical unit may be considered disproportionate. As the base population is decreased by moving to smaller geographies or sub-populations, it becomes easier to find units and discover information. An intruder would also be more likely to have greater confidence in any claim they might make.

Although the local sources reveal the identity of the person, it is the statistics that cause the motivated intruder to start attempting to reveal what is disclosive. Official statistics should not reveal the identity of any respondents with the risk of disclosure. This includes consideration of other relevant sources of information. These sources may be private or public but their relevance is determined by whether they could reasonably to be used to identify a person and reveal information about them. You do not need to consider all local sources, but you should consider information likely to be available to third parties.

An intruder with a special interest in conception statistics discovers from Table A2 that a small number of twins have been born to mothers in specified age groups in England and Wales. The small number in the cell does not lead to direct identification. However, it may prompt an intruder to use knowledge of the age of an acquaintance giving birth to find out she had given birth to a stillborn twin and a living child.

Table A2: Number of twins by age of mother in England and Wales (adapted from Birth Statistics FM1 Table 6.4)

This table uses the following abbreviations:

  • “LM” means “liveborn male”
  • “SM” means “stillborn male”
  • “LF” means “liveborn female”
  • “SF” means “stillborn female”

The first row gives the age of the mother when they gave birth to their twins.

Under 20 20 to 24 25 to 29 30 to 34 35 to 39 40 to 44 45 and over Total
2 LM 93 356 678 982 345 98 921 3473
1 LM and 1 LF 32 234 589 1001 612 121 34 2623
2 LF 78 345 865 943 532 24 103 2890
1 LM and 1 SM 3 10 12 11 15 3 1 55
1 LM and 1 SF 1 3 3 5 3 0 1 16
1 LF and 1 SM 3 1 0 8 3 2 0 17
1 LF and 1 SF 0 2 4 3 4 4 2 19
2 SM 0 3 2 3 0 0 0 8
1 SM and 1 SF 1 1 0 0 1 1 2 6
2 SF 0 1 6 2 3 0 1 13
Total 211 956 2159 2958 1518 253 1065 9120

There are many cells in this table with low counts relating to sensitive information. A motivated intruder having been told informally that an acquaintance aged 45 and over was pregnant with twins could later investigate further to discover the outcome. This could occur if the intruder was aware that the mother later only had one female child. The data reveals that the stillborn child must have been female too. There is no public register of stillbirths, so the information is not in the public domain.

These low counts, including zeros, require protection. Table A2a shows a combination of combining age categories and suppression.

Table A2a: Number of twins by age of mother in England and Wales (adapted from Birth Statistics FM1 Table 6.4)

This table uses the following abbreviations:

  • “LM” means “liveborn male”
  • “SM” means “stillborn male”
  • “LF” means “liveborn female”
  • “SF” means “stillborn female”

The first row gives the age of the mother when they gave birth to their twins.

Under 25 25 to 34 35 and over Total
2 LM 449 1680 1364 3473
1 LM and 1 LF 266 1590 767 2623
2 LF 423 1808 659 2890
1 LM and 1 SM 13 23 19 55
1 LM and 1 SF 4 8 4 16
1 LF and 1 SM 4 8 5 17
1 LF and 1 SF 2 7 10 19
2 SM 3 5 0 8
1 SM and 1 SF c c 4 6
2 SF c c 4 13
Total 1167 5117 2836 9120

Table A2a still contains useful data, although some of the finer detail has been lost. It would need to be assessed to determine whether this table would be sufficient for research purposes.

If detail on stillbirths was not required, the data could be released as in Table A2b, although suppression may be required to protect the cells with a value of 1. Given the table is at England and Wales level, the only risk is likely to be self-identification rather than any attribute disclosure. However, the high sensitivity of the table may persuade a data provider to apply protection. Risk categories of particular variables are discussed earlier in this guidance.

In summary, a table such as Table A2b is likely to be considered sensitive, as any personal information obtained from this table could have an extreme effect on the individual concerned. This means it is less likely to be released as it stands.

Due to the large number of small cells, rounding was not considered for Table A2. This is because it would result in a large proportion of cells being rounded to zero.

Table A2b: Number of twins by age of mother in England and Wales (adapted from Birth Statistics FM1 Table 6.4)

This table uses the following abbreviations:

  • “LM” means “liveborn male”
  • “SM” means “stillborn male”
  • “LF” means “liveborn female”
  • “SF” means “stillborn female”

The first row gives the age of the mother when they gave birth to their twins.

Under 20 20 to 24 25 to 29 30 to 34 35 to 39 40 to 44 45 and over Total
2 LM 93 356 678 982 345 98 921 3473
1 LM and 1 LF 32 234 589 1001 612 121 34 2623
2 LF 78 345 865 943 532 24 103 2890
1 living and 1 stillborn 7 16 19 27 25 9 4 107
2 stillborn 1 5 8 5 4 1 3 27
Total 211 956 2159 2958 1518 253 1065 9120

Identification and self-identification

Where a cell has a large value, risks arising from identification are not usually significant. More consideration is needed where a cell has a small value, particularly if the count is 1. This is because identification or self-identification can lead to the discovery of rareness, or even uniqueness, in the population being considered. For certain types of information – especially information of a sensitive nature – rareness or uniqueness may encourage others to seek out the person in the data. The threat or reality of this could:

  • cause harm or distress to the affected person
  • lead the affected person to claim that the statistics are inadequate to protect themselves or others

The same is true of cells with a value of 2 representing two units, where one of the units contributing to the cell may identify the other. This could occur when groups of people or organisations with similar characteristics who know enough to identify each other appear in the same table. An example could be people from the same household.

To protect against unique identification or self-identification, you should consider – at a minimum – all cells of size 1 or 2 to be unsafe at all but the highest geographical levels. Although direct identification or self-identification may not be a significant risk, protection is often required since identification can lead to attribute disclosure when more than one table is disseminated from a data source. The identified person in an internal cell of a table can become a marginal cell in another table and a new attribute could be learned.

There are no formal rules to follow when publishing tables but if these disclosure risks are considered, the resulting outputs should be of both useful and low risk.

There are some occasions where specific outputs will need to be assessed differently to determine the disclosure risk.

This case study shows an example of identification or self-identification. This will potentially occur from any cells with a count of 1, representing one statistical unit. The same is true of cells with a value of 2 representing two units, where one of the units contributing to the cell may identify the other. This could occur when groups of people or organisations with similar characteristics who know enough to identify each other appear in the same table. An example could be people from the same household.

A table of Live Births outside marriage or civil partnership in geographical area below Clinical Commissioning Group shows a count of 1 when the age groups of both parents are shown. The couple can see that they are the only combinations of these age groups to have had a child in this year. This knowledge could make the couple feel especially vulnerable. An example is seen in Table A3.

Table A3: Number of births by age of parents 

The first row of this table gives the age of the mother when they gave birth to their child. The first column gives the age of the father when their child was born.

Under 20 20 to 24 25 to 29 30 to 34 35 to 39 40 to 44 45 and over Total
Under 20 634 145 65 18 9 2 0 874
20 to 24 1326 1452 764 346 89 23 1 4001
25 to 29 314 1264 1045 890 649 257 26 4445
30 to 34 62 746 784 651 451 365 38 3097
35 to 39 21 259 476 328 214 169 17 1484
40 to 44 13 214 331 267 135 83 15 1058
45 to 49 6 46 54 41 36 23 8 214
50 to 54 1 11 21 32 17 11 6 99
55 to 59 1 3 6 12 8 6 4 40
60 to 64 0 2 3 4 3 4 1 17
Over 64 0 0 1 1 0 1 0 3
Total 2378 4142 3550 2590 1611 944 116 15331

An example of disclosure for this table could occur if an intruder attempts to identify the parents of a child where there is a large gap between the ages of the father and mother, such as the count of 1 for a father aged over 65 and the mother aged 25 to 29.

As with the previous case studies, the table can be protected by collapsing categories, as can be seen in Table A3a.

Table A3a: Number of births by age of parents 

The first row of this table gives the age of the mother when they gave birth to their child. The first column gives the age of the father when their child was born.

Under 20 20 to 24 25 to 29 30 to 34 35 to 39 Over 39 Total
Under 20 634 145 65 18 9 2 874
20 to 24 1326 1452 764 346 89 23 4001
25 to 29 314 1264 1045 890 649 257 4445
30 to 34 62 746 784 651 451 365 3097
35 to 39 21 259 476 328 214 169 1484
40 to 44 13 214 331 267 135 83 1058
45 to 49 6 46 54 41 36 23 214
Over 49 2 16 31 49 28 33 159
Total 2378 4142 3550 2590 1611 1060 15331

If the cells of size 2 are considered disclosive then suppression could be applied. Another option would be to collapse the table further and create categories 45 and over (for the father) and 35 and over (for the mother).

Rounding to base 3 could be an option here, although the data would become much less useful as many cells have low counts. This can be seen in Table A3b. Simple rounding is applied where each cell is rounded to the nearest multiple of 3. This often results in the table being non-additive.

There is a question about what to do with cells which have a count of zero in the original table. Should they be represented by dashes or zeros? If there is a distinction between the two it will be easier for an intruder to unpick the table, as it would be obvious that cells with counts of zero had been rounded down from a larger number. However, it does increase the level of detail published. A decision will need to be made on what to publish.

Table A3b demonstrates a distinction between zeros. Both dashes and zeros are displayed to highlight empty cells in the unrounded table.

For some outputs it may be that ‘0’ is appropriate when the figure is a true zero, and not when a figure has been rounded to zero. In this case, ‘[low]’ should be used for non-zero low counts. As discussed, this will affect the confidentiality of the table. Better practice would be not to distinguish between ‘real’ zeros and those that have been rounded to zero. An exception can be made for structural zeros (or cells for which an event cannot occur, such as 0 to 5-year-olds being mothers).

Table A3b: Number of births by age of parents 

The first row of this table gives the age of the mother when they gave birth to their child. The first column gives the age of the father when their child was born.

Under 20 20 to 24 25 to 29 30 to 34 35 to 39 40 to 44 45 and over Total
Under 20 633 144 66 18 9 3 - 873
20 to 24 1327 1452 765 345 90 24 0 4002
25 to 29 315 1263 1044 891 648 258 27 4446
30 to 34 63 747 783 651 450 366 39 3096
35 to 39 21 258 477 327 213 168 18 1485
40 to 44 12 213 330 267 135 81 15 1059
45 to 49 6 45 54 42 36 24 9 213
50 to 54 0 12 21 33 18 12 6 99
55 to 59 0 3 6 12 9 6 3 39
60 to 64 0 3 3 3 3 3 0 18
Over 64 0 0 3 1 0 0 0 3
Total 2379 4143 3549 2589 1611 945 117 15330

Mosaic, or jigsaw effect

A more detailed scenario to consider is that of the mosaic, or jigsaw effect. This is where different data releases – each of which may be deemed safe in isolation – can be pieced together to effect disclosure. Another example is where there are multiple releases from a similar collection over time, perhaps for successive months or quarters. Change between each release may be small, which may make it easier to identify people in the data.

As we see a growth in publicly available sources, including social media, this risk has increased considerably.

In this example, the hypothetical administrative dataset of 12 variables comprises:

  • four sensitive variables of high impact – the variables are limiting long term illness, sexual orientation, income, and religion
  • four visible variables which can often be determined by observation – these variables are sex, (broad) ethnic group, household composition, and occupation
  • four other variables – these variables are country of birth, age, number of children, and level of education

Various tables are formed and released from these data.

The risk is that an intruder could use the visible variables to identify a person and find out a piece of information relating to a sensitive variable. The other variables can help as additional matching variables or catalysts.

There are 4 releases from these data.

Release 1 contains data about sex, ethnic group, country of birth, age, number of children, and level of education. This release may be safe. The visible variables are sex and ethnic group, and these may not be enough to identify a person.

Release 2 contains data about limiting long term illness, household composition, country of birth, and age. This release may be safe. The visible variable is household composition, which may not be enough to identify a person.

Release 3 contains data about sexual orientation, occupation, number of children, religion, and level of education. This release may be safe. The visible variable is occupation, which may not be enough to identify a person.

Release 4 contains data about limiting long term illness, sexual orientation, country of birth, and religion. This release contains a lot of sensitive information but does not contain any visible variables, which should make an intruder unable to identify a particular person in the data.

While these releases may be safe in isolation, an intruder could layer them to help identify a person in the data. Releases 1 and 3 both contain data about level of education and number of children, which could reveal other attributes when considered together. Other releases also contain common variables:

  • releases 2 and 4 both contain data about limiting long term illness and country of birth
  • released 3 and 4 both contain data about sexual orientation and religion
  • releases 1 and 2 both contain data about country of birth

These releases have the potential to reveal other attributes when they are considered together. It would not be easy for an intruder to garner this information unless combinations of these variables approach uniqueness, but it is possible to see how the releases could be used to identify a person.

It is worth remembering that linking and disclosure may only be possible for one person, or a small number of people. But that may be the person with the most unusual combination of characteristics, which could make them the most vulnerable.

This mosaic effect requires considerable thought and each combination of outputs needs to be considered as a whole. The ultimate aim is to be pragmatic and not to think that the worst case will always happen. As always, the aim is to balance the disclosure risk of releasing a group of linked tables with the need to ensure they remain useful.

Assessing the risk

Start by defining the variables carefully. How sensitive are the variables? What would be the impact of disclosure?

Then, think about the order of release. Are the tables released together or at different times? Are they ad-hoc requests to the same customer? This is where a log of releases would be recommended. A log of releases would allow data analysts to have an easily accessible record of previous releases so they could see which combinations of variables (and categories) have been published previously or made available to the same user. Consider whether a new publication could cause disclosure issues. Remember that releases to the same customer would be likely to be more of a problem although the legal position is the same as releasing to different customers.

Consider whether there are other datasets available in the public domain with some of the same variables. They may be collected at different times, through different modes and have different measurements, but will still add a little to the overall disclosure risk.

Creating and releasing tables

If a defined set of tables has been released previously, users are likely to expect they will be published to the same specification in future releases. Users may complain if the publishing strategy is changed, so you may want to consider this when you are selecting the disclosure control method.

If many tables are to be released from a particular dataset, it may be advisable to protect the underlying microdata before creating your tables. This helps address any disclosure by differencing. Read more about protecting microdata created from social surveys.

Flexible table builders are becoming more of an option for users who want to create their own tables. Users can use them to download many tables from the same microdata.

One example is the create a custom dataset (CACD) tool developed for Census 2021 in England and Wales.

There are multiple disclosure issues with flexible table builders. The main challenges are:

  • univariate uniques — this is where there are single counts in a requested table with only one variable
  • differencing
  • ensuring disclosure risk has been reduced to an acceptable level
  • how a flexible dissemination system affects microdata releases

Read Stephanie Blanchard’s paper about the methodological challenges of protecting outputs from a flexible dissemination system, published in the Survey Methodology Bulletin (number 79)

Population at risk assessments

The effect of geography on disclosure risk has already been discussed in this guidance. This can be expanded to consider risk for different population sizes alongside variables of differing sensitivity.

You should also consider the safety of frequent releases of the same measures, especially those that are supplied daily rather than monthly, quarterly, or annually. There are likely to be small fluctuations between each release, which means the disclosure risk could be high – even where the population at risk may also be high.

You should carry out a risk assessment exercise to develop suitable confidentiality rules for different datasets. In practice, it is likely that producers of statistics will find that risk levels for outputs from the administrative data can be defined by the population at risk with a further breakdown by the impact of any identification. This means considering whether the output contains any high, medium, or low impact variables.

Decisions on the likelihood and impact of identification should be made by someone with detailed knowledge of the data. When deciding on the risk level and level of protection required, thought will need to be given both to the:

  • likelihood of an identification (based on population at risk)
  • impact that an identification would have

Tables with sensitive defining variables will almost certainly have a greater impact if an individual is identified correctly.

Impact through the release of specific combinations of variables (or ‘statistical impact’ as described earlier in this guidance) is a subjective term. However, possible categorisations for a small number of common variables are given in this section. Some of these are the same as the broader selection of variables defined as sensitive by the DPA.

Variables with higher impact are more sensitive, meaning that any disclosure would cause a great deal of distress to the person concerned.

High impact variables

These may include:

  • income or wealth — any financial variables are generally considered to be high impact
  • ethnic origin
  • religious beliefs
  • physical or mental health
  • some types of crime — these variables are high impact for victims
  • sexual orientation
  • gender identity
  • economic activity — especially responses of ‘unemployed’ or ‘never worked’
  • industrial classification — this variable is considered to be high impact at the most detailed level
  • occupation
  • qualifications — especially level of qualification and details of the subject studied

The lower the level of geography at which a table is released, the greater the level of risk for these variables. At local area level (such as Census Output Area), it is possible that variable combinations, including one or more of the variables in the ‘high impact’ list, will be unique for an individual.

Medium or low impact variables

These may include:

  • marital status — where this is not covered by vital registration
  • single year of age
  • household composition
  • geography — at Local Authority, or equivalent
  • industrial classification — where this is grouped into a small number of categories
  • sex
  • age group
  • house type

The risk will be lower if the geography is at region level or above.

As previously stated, counts of 1 or 2 for tables at a relatively low level of geography need to be thought of as possibly disclosive. However, as populations can differ widely across administrative geographies, the population at risk for a table produced from a lower level of geography will not be consistent. It is difficult to give exact guidance, but one approach is to consider ‘population at risk’. This is not always easy to define conclusively but this section gives examples of population sizes. Please note these examples are indicative rather than prescriptive. The relationship between risk and impact is also considered for both high impact, and medium to low impact.

This is equivalent to the population of a small Local Authority or Unitary Authority with only a few being smaller than this. The term Local Authority also covers Unitary Authority for the remainder of this section.

In a population of this size or larger, any attempt at identification by an intruder would almost certainly not unsuccessful.

Tables of low counts (1s and 2s in the margins) where one or more variables are high impact could be examined for disclosure issues. But in most cases, tables based on this population at risk can be released with no application of disclosure control.

If there is a high impact variable in the table then you could consider protecting counts of 1 or 2, though in many cases there will be no problems. If there are no high impact variables, you can publish without additional protection.

A similar methodological approach to that in the previous section can be followed.

As the population is smaller, tables with high impact variables will possibly require disclosure control to be applied for tables with counts of 1 or 2, although each case should be judged separately.

For a concentrated population of this size in an urban area an intruder could possibly identify a friend or relation in a table, although some work would be required to do this with great confidence.

There is some risk for tables with at least one high impact variable, so disclosure control may be appropriate for cells of size 1 or 2 in the body of the table and the margins. Rows or columns dominated by zeros should also be checked. Where no variables are high impact, disclosure control may not be necessary, although low counts especially in the margins should be looked at closely.

At this population size it becomes easier for an intruder to discover a friend or neighbour in a table. There are a limited number of people who could be in the table, so it might only take a neighbour with a high level of curiosity a short amount of time to make a correct identification with great confidence.

If there is at least one high impact variable in the table, disclosure control will be required for tables with cells of size 1 or 2 in the body of the table or the margins. Rows and columns dominated by zeros will also require the application of disclosure control. If there are no high impact variables, low cell counts may require protection, but each table will need to be looked at individually.

An intruder may be successful in finding particular people in tables based on a population this low. They are also likely to be very confident in their claim.

Disclosure control should be applied to these tables whether the variables are of high or low impact. Any cells of size 1 or 2 and row and columns dominated by zeros will require protection.

The levels of geography at which data are likely to be released (such as LAD) vary considerably in population size. This is why population at risk is used alongside the impact of a variable to define the risk of a table.

The values for population at risk are useful as a guide and suggest levels at which an intruder might be more inclined to try to identify an individual in a published table. Smaller populations at risk require more protection. However, experts in different areas of official statistics may want to choose different values for population at risk while maintaining the structure outlined in this guidance.

The likelihood of an attempt at identification and its impact may be heightened and additional protection required if:

  • any other disclosive situations are likely to occur
  • statistical units are represented more than once in the table — for example, if the statistical unit is a patient and the table reports annual hospital admissions, then a cell of 4 could represent the 4 times the same patient was admitted
  • groups of statistical units are represented in the table — for example, people from a particular household
  • tables based on the dataset have already been released — the likelihood of identification may increase due to linking and differencing with these past releases, which means protecting large databases may be important
  • other freely available datasets can be linked to the tables
Back to top of page

Legal, ethical and trust considerations

The 2018 Data Protection Act and Code of Practice for Statistics have already been discussed in this guidance. There is further discussion of legal, ethical and trust considerations in this section.

When establishing whether confidentiality protection is required for a particular statistic, it is necessary to consider:

  • public trust and cooperation
  • legal rights and obligations
  • national and international standards for statistics

This will help determine acceptable disclosure risks and unacceptable disclosure risks.

The production and use of statistics depends on the co-operation and trust of citizens. Such trust cannot be maintained unless the privacy of individuals’ information is protected – and perceived to be protected. Failure to respect privacy might result in harm or distress to a specific individual. Sensitive personal records, therefore, need to be strictly confidential. On the other hand, there is a legitimate public interest in having ready access to aggregate or summary statistical information.

The legal framework covering the use of personal information as defined by the SRSA and DPA is complex. When such information is transformed into statistics, the legal framework is much simpler. The statistical information can be widely and freely used, provided confidentiality protection has been applied such that it is no longer likely that the information can be related to specific identifiable individuals.

When the information in a statistic does not relate to an identifiable individual (either on its own or in combination with other information likely to be available), there can be no breach of the duty of confidence owed or any conflict with data protection or human rights legislation.

Information available in the public domain is not necessarily risk-free when presented in a table or as another statistic. Statistical disclosure control methods may modify the data or the design of the statistics, or a combination of both. They will be judged sufficient when the guarantee of confidentiality can be maintained.

Statisticians or researchers should have no interest in the individual statistical unit, other than to distinguish one unit from another for statistical purposes. This could include conducting data matching or linking exercises for statistical purposes, where identified data are essential for quality reasons.

In contrast, an intruder is someone who, for whatever reason, wishes to distinguish one statistical unit, so they can treat that unit separately or differently from the other statistical units in the dataset for a non-statistical purpose. Producers of statistics ought to be aware that intruders may use other public and private sources of information in their attempt to identify an individual or other member of the published table.

Phrases related to the word ‘identity’ are used frequently in legislation concerning disclosure control and data confidentiality. The 2018 Data Protection Act defines an “identifiable living individual” as a living person who can be identified, directly or indirectly, in particular by reference to both direct identifiers and factors specific to the physical, physiological, genetic, mental, economic, cultural or social identity of the individual.

Data relating to an identifiable living individual are defined as personal data. The Act states that it is unlawful to obtain personal data without consent of the data controller.

Back to top of page

Selecting disclosure control methods

Where required, disclosure control methods can be used to reduce the risk by modifying the unsafe cells. The choice of method must balance the uses to be made of the information and simplicity of approach. The method chosen should ensure that the released data is of high utility and disclosure control is applied appropriately. The data should not be over protected or under protected.

Table redesign

This is recommended as a simple method that will minimise the number of unsafe cells and preserve original counts. However, the use of this method should be balanced against consistency in publication plans when initially planning the design of tables.

You can use table design to disguise unsafe cells by:

  • grouping categories within a table
  • aggregating to a higher level geography or for a larger population sub-group
  • aggregating tables across several years, months, or quarters

The advantages of this approach are as follows:

  • the original counts in the data are not damaged
  • it is easy to implement

The disadvantages of this approach are as follows:

  • the detail in the table will be reduced
  • you may also be restricted in your redesign if there are policy or practical reasons that require a particular table design

Examples

Examples include:

If unsafe cells remain in the output tables, further protection methods should be considered to disguise them. If table redesign is not a feasible solution, the recommended method for post-tabular protection for most frequency tables is controlled rounding (which preserves the additivity of the table). However, this method requires specialist software and a linear programming solver and therefore will not always be practical. In some cases, if the number of unsafe cells is low then cell suppression can be an alternative method. Controlled rounding and cell suppression can be implemented in specialist software such as Tau-Argus.

Cell modification methods

This involves changing the cell values. Some or all of the values are different in the published data compared to the original. Methods are cell suppression, addition of noise and rounding.

When cells are suppressed, unsafe cells are not published. They are suppressed and replaced by a special character, such as ‘c’, to indicate a suppressed value. Such suppressions are called ‘primary suppressions’. To make sure that the primary suppressions cannot be derived by subtraction from totals, it may be necessary to select additional ‘safe’ cells for secondary suppression.

Advantages

The advantages of this approach are as follows:

  • original counts in the data that are not suppressed are not adjusted
  • cell suppression can provide protection for zeros
Disadvantages

The disadvantages of this approach are as follows:

  • most of the information about suppressed cells will be lost
  • secondary suppressions will hide information in safe cells
  • information loss will be high if more than a few suppressions are required
  • any disclosive zeros will need to be suppressed to reduce the disclosure risk
  • this method does not protect against disclosure by differencing
  • this method can be complex to implement optimally if more than a few suppressions are required and is particularly complex for linked tables
Example

Cell suppression is used in detailed characteristics of birth tables. In the birth characteristics dataset, rates are not displayed when there are fewer than 3 births in a cell.

This is where counts in tables can be adjusted, or ‘perturbed’, up or down by small amounts.

Advantages

The advantages of this approach are as follows:

  • adding noise provides uncertainty to any disclosure — especially if some non-structural zeros are perturbed
  • adding noise can provide significant protection against disclosure by differencing if each table is perturbed independently
  • the cell key method used in Census 2021 maintains consistency where the same cell appears in different tables
Disadvantages

This method can affect the usefulness of the table where cells and tables are perturbed independently.

Examples

Census 2021 tables are subject to the cell key method, which:

  • adds noise to every cell — some noise will be ‘0’
  • maintains consistency of cells in different tables
  • perturbs a small number of zeros

Some breakdowns of variables will lead to inconsistencies in totals, but this method does allow every cell to have a value.

This method involves adjusting the values in all cells in a table to a specified base. This creates uncertainty about the real value for any cell while adding a small but acceptable amount of distortion to the data.

Advantages

The advantages of this approach are as follows:

  • counts are provided for all cells
  • rounding provides protection for zeros
  • rounding protects against disclosure by differencing and across linked tables
  • controlled rounding preserves the additivity of the table and can be applied to hierarchical data
Disadvantages

The disadvantages of this approach are as follows:

  • rounding cannot be used to protect cells that are determined unsafe by a rule based on the number of statistical units contributing to a cell — for example, if a cell had an original count of 17 events all associated with one practitioner, then rounding this to 15 means that the count still relates to only one practitioner and the unsafe cell is not disguised
  • random rounding requires auditing
  • controlled rounding requires specialist software and a linear programming solver, leading to the method not always being a practical solution
Example

Counts from the New Zealand Census are rounded to base 3 in their outputs.

How to handle zeros when rounding

When rounding is implemented, there is a question about what to do with cells which have a count of zero in the original table. Should they be represented by dashes or zeros? If there is a distinction between the two it will be easier for an intruder to unpick the table as it would be obvious that cells with counts of zero had been rounded down from a larger number. However, it does increase the level of detail published.

For some outputs it may be that ‘0’ is appropriate when the figure is a true zero, and not when a figure has been rounded to zero. In this case, ‘[low]’ should be used for non-zero low counts. But you should think carefully about the confidentiality implications of doing this. Better practice would normally be not to distinguish between ‘real’ zeros and those that have been rounded to zero. An exception to this approach can be made for structural zeros (cells for which an event cannot occur, such as children between the ages of 0 and 5 being mothers).

Database modification methods

These methods involve adjusting the underlying microdata prior to tabulation. This is possible where a data provider has access to the individual record level data.

This involves swapping pairs of records within a micro-dataset that are partially matched to alter the geographic locations attached to the records, leaving all other aspects unchanged.

Advantages

This method can – to some extent – protect against disclosure by differencing. The advantages of this method are as follows:

  • only needs to be applied once to the microdata
  • allows for flexible table generation
  • can target risky records
  • gives consistent and additive tables
  • leaves counts at high geographies unaffected
Disadvantages

The disadvantages of this method are as follows:

  • lots of swapping may be required to disguise all unsafe cells
  • distributions in the data at low geographies will be distorted
  • this method is not transparent to users — it may appear as if disclosure control has not been carried out, and more metadata may be needed to support the use of this method
  • record swapping may lead to a perceived risk of disclosure, as cells with low counts will be published
  • the theory and practicality of this method may not easily be understood — this means that considerable communication and education will be required alongside outputs
  • calculations relating to loss of data utility and doubt may need to be completed before all output tables are produced — other tabulation methods may be required if risks remain
Example

Record swapping has been used as the primary method of confidentiality protection for the UK censuses in 2001, 2011 and 2021.

A small number of records may be unique in the data for several variables. Rather than protecting tables using these variables, it can be simpler to remove the record.

Advantages

Less protection is required in the published tables without having to allow for an outlying record.

Disadvantages

The disadvantages of this approach are as follows:

  • this method involves making a subjective decision to remove information from the dataset
  • users of the data may not know this has taken place or the methodology behind the removal of certain records
Example

This is more likely to be a technique used before microdata are released. There are no publicly available examples of the use of this method for ONS tabular data.

There are many other methods of disclosure control.

For example, an alternative approach would be to create a synthetic dataset which maintains all the properties of and relationships in the true dataset. From these data, non-disclosive tables can be created.

Equally, alternative methods for presenting data can be considered to provide users access to information without disclosing the underlying data. In many cases, this will provide a more robust analysis than reliance on the accuracy of small cell counts. These could include presenting data graphically with limited detail in scale or providing commentaries or analytical outputs.

Back to top of page

Implementing your selected method, or methods

This guidance will allow data providers to set disclosure control rules and select appropriate disclosure control methods to protect different types of published tables of statistics based on administrative sources. The most important consideration is maintaining confidentiality, but these decisions will also accommodate the need for clear, consistent and practical solutions that can be implemented within a reasonable time and using available resources. The methods used will balance the loss of information against the likelihood of individuals’ information being disclosed. Data providers should be open and transparent in this process. All decisions should be documented, along with the whole risk assessment process so these can be reviewed.

You should think about the relationship between risk and usefulness when setting disclosure control rules for tables produced from administrative data. It is often impractical to aim for zero risk as the resulting outputs would be of little practical use to prospective users of the data. This means that any released data will have a small level of associated risk, but there will also be sufficient uncertainty that any attempted identification or attribute disclosure would be correct.

There is also a relationship between disclosure risk and disclosure impact. What is considered a sufficient level of uncertainty will differ depending on the sensitivity of the release and the impact of a disclosure. This is difficult to quantify, and decisions ought to be taken by those most familiar with the data. A higher level of uncertainty would be required when the table involves a sensitive variable or variables, typically those of high impact. Disclosure of any sensitive or high impact information could have a significant and longstanding effect upon both the affected person and the organisation.

The effect of a false claim must also be considered. For example, a false claim that relates to a person and their involvement in criminal activity could cause great harm or distress, despite being incorrect.

When looking at published statistics, users should be aware that:

  • the dataset has been assessed for disclosure risk
  • methods of protection may have been applied

For quality purposes, users of a dataset will be provided with an indication of the nature and extent of any modification due to the application of disclosure control methods.

Any techniques used may be specified, but the level of detail made available should not be enough to allow the user to recover disclosive cell counts.

Data providers may need to make judgements in a wider context than the specific statistics that they are producing at any one time. They need to be aware of decisions made by others within their organisation, either in the past or for similar sectors. This helps ensure decisions are consistent with the policy and strategic position of the organisation. Decisions also need to be made in the context of wider information governance arrangements both in an organisation and more widely.

Data producers should also be prepared to respond to claims of disclosure, whether these claims are correct or incorrect.

When providers share data with a second party to be published (assuming this complies with any legal or policy requirements), they must ensure the second party follows the general guidance and any specific confidentiality rules that have been developed. These should be stated in the data sharing agreements. This will ensure consistency between published statistics derived from the same source.

Any change in disclosure control rules for a published statistic raises the issue of revisions to previous releases. In general, new disclosure control rules will be implemented for future releases and the rules will not apply to past releases. An exception can be made in cases where the disclosure control rules are altered to allow more data to be released. In these cases it may be feasible to re-release older datasets with greater detail.

Statistical confidentiality is a public interest which will normally outweigh other relevant public interest in disclosing the underlying confidential records to the public. There will be rare occasions when the public interest in disclosing the records might outweigh the public interest in confidentiality. An example would be a requirement to publish details relating to a spate of crimes in a local area. Such decisions must only be taken at the highest level and in consultation with the Head of Profession. In many cases it will be found that these records are not statistics but factual information, which is subject to a different set of rules or guidance. The difference between statistics and factual information is discussed in the FOI Act.

Freedom of information requests

Under the Freedom of Information (FOI) Act ‘statistical information’ and ‘factual information’ are treated differently within the section 35 exemption. The equivalent exemption in the Freedom of Information (Scotland) Act 2002 (FOISA) is section 29. You can find guidance on the FOI exemption on the Information Commissioner’s website and guidance on the Scottish equivalent of the FOI exemption on the Scottish Information Commissioner’s website.

Put simply, factual information is the record of events or administrative actions. Statistical information is the outcome of a transformation, aggregation or analysis of such records performed using a repeatable methodology. Records of a disease, for example, are factual information, and the aggregation and analysis of those records is statistical information.

People have a general right of access to information on themselves held by public authorities through the FOI or FOISA, as long as this does not contravene confidentiality constraints. Confidentiality policy developed using this guidance can be used to help decide which exemptions in the Act are relevant, and which should be cited when withholding confidential statistical information. It is good practice to explain a general policy for the withholding of information, but this must be done in addition to – and not in place of – the exemptions in the FOI Act.

FOI requests should always be considered on a case-by-case basis. There may be cases when decisions about a case appear to be inconsistent with the general policy for the publication of statistics. This does not mean that the policy is wrong, since it has been developed for use in a production process rather than assessment of every individual cell count. Whilst confidentiality must always be maintained, a decision made under FOI to provide information in a form different to the published outputs is compatible with this guidance.

Back to top of page

Help and support

If any problems arise when applying statistical disclosure control to tabular outputs from administrative data, please contact the SDC Expert Team by emailing SDC.Queries@ons.gov.uk.

See a flowchart of the SDC process from receiving data to output.

Back to top of page