The scope of this guidance
This guidance:
- covers how to assess microdata risk and methods of mitigating this risk
- considers suitable methods of disclosure control required to ensure the confidentiality of released microdata
- highlights specific issues, such as those relating to longitudinal data and large households
It also gives general guidance for ensuring that microdata derived from social surveys adhere to the legal restrictions necessary for both:
- personal data as defined by the Data Protection Act (DPA) 2018
- personal information as defined by the Statistics and Registration Service Act (SRSA) 2007 for data releases from the Office for National Statistics (ONS)
It is important to note that the Data Protection Act (DPA) must be followed for all public data releases, taking particular care of any sensitive personal variables.
This guidance can be read alongside advice on the production of an open data set.
Four case studies are included in this guidance. They are added at relevant points in the text. These case studies all follow a similar structure with common themes and can be utilised independently of the associated text.
Who this guidance is for
This guidance will be most useful to those who are preparing safeguarded files to be released under an End User Licence (EUL), possibly through the UK Data Service or other research environments.
Access levels for microdata
Microdata can be released under different access levels. There are three main levels.
Public Use Files (PUF)
There is negligible disclosure risk for these files. Considerable disclosure control may have already been applied and they are of most use as training or teaching datasets.
Safeguarded files
These are usually released under an End User Licence (EUL) with low (but not negligible) disclosure risk. These can be used for general research purposes.
These datasets are not personal information under SRSA but may be personal data under DPA. This is because identification may be possible using private unpublished sources. Users must sign a declaration before gaining access to the data. One condition of an EUL is that the user must promise to maintain the confidentiality of the data and should not attempt to identify any person, household or business in the data, nor to claim to have made an identification (in the case of spontaneous recognition).
Disclosive microdata
These microdata are only available to Approved or Accredited Researchers (AR), usually in a secure setting. Accredited Researcher is the term used at the Office for National Statistics (ONS), but other departments may use different terminology. Outputs are checked before they are released with little disclosure control applied to the microdata. Only minimal disclosure control is required for data in a secure setting, so this level of access is not discussed further in this guidance. These datasets are used for the most detailed analysis.
If you are using microdata in the background to populate interactive tabulation tools (known as table builders), you should also include protection against disclosure. This may take the form of both pre-tabulation and post-tabulation methods, which may include:
- applying protection directly to the microdata
- some form of table design
- altering, or ‘perturbing’ some cell counts
For further advice, please contact the Statistical Disclosure Control Expert Group by emailing SDC.Queries@ons.gov.uk.
Risk assessment for microdata from surveys
When assessing the disclosure risk of microdata, you should first consider the most likely intruder scenarios. These are assumptions about:
- what an intruder might know about respondents
- what information will be available to match against the microdata and potentially make an identification
The exercise of considering scenarios can help indicate some of the variables which are most likely to be used by an intruder. Some scenarios may be specific to the dataset.
Scenarios to consider
The definitions of scenarios and the corresponding key variables were initially developed by Angela Dale and Mark Elliot. You can read more about this in their papers:
- “Disclosure risk for microdata Work package DM 1.1 What is a key variable?“
- “Scenarios of attack: the data intruder’s perspective on statistical disclosure risk“
This was later expanded in the paper “Formalizing the Selection of Key Variables in Disclosure Risk Scenarios” by Elliot, Mackey and Purdham (2011).
There has also been considerable research into quantifying the level of risk, an example being Natalie Shlomo’s paper on “Methods to assess and quantify disclosure risk and information loss under statistical disclosure control“.
Two of the scenarios defined by Elliot, Mackey and Purdam were the commercial database crossmatch and spontaneous recognition. For safeguarded microdata, these are among the main scenarios which need to be considered, and these are described in this guidance.
It is important to note that you may need to consider other scenarios, either because of the particular characteristics of a survey, or if more publicly available data is discovered.
An intruder with access to the safeguarded dataset and the published datasets can use key variables to match these against the microdata.
Examples of such published information include:
- the Electoral Register
- vital registration – live births, deaths, marriages, divorces
- commercial datasets – such as consumer profile databases which may be purchased at a reasonable cost by any member of the public
- information available on social media platforms.
- published tables from the same dataset or other related microdata releases – these could include tables or microdata from previous waves (a distinct survey) of the same series that may have a panel element (meaning that the same people appear in the data in successive waves)
We need to think about how sensitivity can be used in terms of risk assessment.
Sensitivity can be difficult to define in the context of published data. Depending on their circumstances, some people might consider a particular variable to be sensitive where others might not. The DPA and UK General Data Protection Regulation (UK-GDPR) both discuss variables that should be considered especially sensitive and likely to be sensitive for most people, if not all. These variables should be treated more carefully. In more detail, the UK-GDPR defines special category, or ‘sensitive’ data as:
- personal data revealing racial or ethnic origin
- personal data revealing political opinions
- personal data revealing religious or philosophical beliefs
- personal data revealing trade union membership
- genetic data
- biometric data (where used for identification purposes)
- data concerning health
- data concerning a person’s sex life
- data concerning a person’s sexual orientation
An income variable does not appear on the UK-GDPR list, but it may well be considered sensitive by many members of the public. Other variables could be sensitive for a small group of people. These might include Occupation, Industry, and Qualifications. Other variables may be very sensitive for a small number of people in specific categories, such as Sexual Orientation and Gender Identity (SOGI). Sensitivity is therefore contextual at an individual level.
Scenario 1 assumes that an intruder could attempt to link published information with safeguarded microdata using an identification key comprising demographic variables. They would then have the direct identifiers of name and approximate location, linked with all the other information in the safeguarded dataset. There is no guarantee of these being correct matches, although many records could be correctly matched.
The perception of a match being correct is also important. Even incorrect matches can have a serious impact on the wrongly identified person. For variables defined as sensitive, this could lead to personal details being linked to a name and location (perhaps an address in certain cases), which could increase the distress caused to the identified person.
An intruder may spontaneously recognise a person in the microdata by means of known information. This information might be published and therefore widely available, or known to a small number of people with knowledge relating to a particular person or group of people. For example, spontaneous recognition can occur when a respondent has unusual characteristics and is either an acquaintance, or a well-known public figure.
The key variables for this scenario include:
- name
- age
- sex
- marital status
- income – in the case of high earners
- occupation – job title, which may be equivalent to 4-digit occupation and industry classifications
- resident geographic level
Other variables may be considered visible for some respondents, such as ethnic group, religion, or accommodation type.
Both scenarios (and especially scenario 2) can have increased likelihood and risk if the intruder has response knowledge. That is, that they have some knowledge that a particular person is a respondent and is included somewhere within the data.
Another likely scenario is where there is a clearly defined population. An example of this could be all patients diagnosed with a very unusual condition. These scenarios can make smaller samples riskier as the known sample member will be more distinct in the data. However, the overall disclosure risk is nearly always greater for larger sample sizes, as described in the upcoming paragraph on sample size.
There are several other factors to consider as part of the risk assessment.
Key variables
The identification of key variables is an important element in the role of an intruder. By considering these scenarios – and others if appropriate – we can identify which variables an intruder is likely to combine to create an identification key. Such a key could be used to attempt to match the microdata with other published data to which the intruder has access. This would enable them to identify one or more records that refer to particular people. The identification risk comes from people within the microdata that are both:
- sample uniques – records that are unique within the sample on a given combination of variables present in the dataset
- population uniques – records that are unique within the population on a given combination of key variables present in the dataset
This is because this increases the probability of a match, and therefore that of re-identification. A sample unique could be presumed to be a population unique by an intruder and in some cases, this could be a correct assumption.
The disclosure risk comes from the other variables in the safeguarded dataset. Provided that the identification risk is reasonably small, the data may not be considered personal information. Historically, the level of protection was deemed acceptable if it would take a disproportionate amount of time, effort and expertise for an intruder to identify a statistical unit, or to reveal information about that unit not already in the public domain. As time has passed, what constitutes ‘disproportionate’ has changed. Twenty years ago, an intruder would have been assessed against a government analyst using standard office tools for a few hours. In the present day, we must consider specialist intruders using specialist tools, with access to the internet and computational power reducing the amount of time needed.
It follows that the variables which are most likely to need protection to prevent identification include demographic indicators, such as:
- geography
- household composition
- ethnicity
- occupation
- sexual identity
Salaries and household income are also key variables. The level of matching between datasets can be improved when income data are used as part of a matching key. Although tax data are not publicly available in the UK, some income-related data are in the public domain. Examples are the salaries of company directors and senior civil servants, along with very high salaries and bonuses, which are all published (such as the list of top earners at the BBC). Some of this information at an aggregated level could be discovered through searching the relevant websites, but access to the microdata would be required to link a salary to a specific individual.
k-anonymity
This is a criterion sometimes used to ensure that there are at least ‘k’ records within a de-identified, or pseudonymised, microdata set that have the same combination of indirect identifiers. It is sometimes called ‘using a threshold of k’ and the values of 3 or 5 are usually used. This is one approach to protecting privacy, ensuring that any given record in the microdata is indistinguishable from at least k-1 other records based on certain identifying variables. If any combination of these variables has fewer than k contributors, categories are grouped or variables are removed from the dataset until k-anonymity is achieved.
The effectiveness of k-anonymity depends largely on the choice of identifying variables and the value of k selected. When choosing k, you should consider any other information in the dataset that could be used to distinguish between people with the same characteristics. It may be necessary to apply further disclosure control methods to the dataset.
The main limitation of the k-anonymity approach is that it can result in high information loss, which affects the usefulness of the data.
The NHS Digital “What About Youth Survey”
This survey was launched in 2015 as part of a government pledge to make improvements to the health of young people. It consisted of a postal survey of 15-year-olds which asked questions about:
- smoking
- drinking
- drug misuse
- diet
- physical activity
- bullying
- wellbeing
The survey was designed to provide comprehensive results at Local Authority (LA) level. The achieved sample size was around 120,000 out of an eligible population of around 300,000.
In most LAs most 15-year-olds were sampled and sent the questionnaire. In a few LAs all 15-year-olds were sampled.
The survey organisers stated that:
“The dataset already has some inherent disclosure control in that even though all children were sampled in some LAs, not all responded. Therefore, finding there is only one Asian girl of a certain religion and you know someone who fits that description, then you can not necessarily deduce how she answered the personal questions about smoking, drug misuse etc. from the dataset as there may be others who fit that description who did not return their questionnaire.”
The aim was to deposit an individual level file for access through the UK Data Service. One possibility was to place the file for re-use under a Special Licence (SL) which restricts access. However, the preference was to make the data available to everyone, as a safeguarded file released under an End User Licence (EUL).
NHS Digital started the disclosure control process by looking at potentially identifying combinations of data items such as:
- LA
- gender
- ethnic group
- religion
- sexual orientation
- family composition
- deprivation quintile
- disability
Many combinations had fewer than 5 contributors. To protect these low counts, it would have been necessary to group or remove variables of key interest. However, the LA code could not be removed as Local Authority estimates were essential. Many of the topics the survey covers differ greatly by Ethnic Group or deprivation, which meant that it was not an option to group these variables.
NHS Digital considered alternative methods, such as adding some random perturbation into the dataset instead. At this point they requested methodological advice from the disclosure control team.
Risk assessment
The proposal to combine categories until there is only a small percentage of cell counts below 5 in a cross-tabulation of the listed variables is a good approach. However, it is likely that this would result in significant loss of detail.
There were several specific risk assessment points.
Preferred release route
The preferred release route was the safeguarded route. The use of private knowledge is not normally considered as a risk scenario here, but one scenario was considered appropriate in this case.
The scenario involves a parent who appropriately obtains access to the data. They may be motivated to see if they can find their teenage son or daughter in the data (though this would normally be a breach of access conditions). In many cases, there would be sufficient variables for this. These include:
- household structure
- year and month of birth
- 18-category ethnic group
- 8-category religion
- whether they have a long-term illness or disability
Issues around low counts
The k-anonymity approach of looking for a minimum number of respondents across identifying combinations of the variables for specific questions is a useful starting point. Possible variables are:
- LA
- Gender
- IMD quintile (derived from postcode)
- Ethnic Group
- A measure of disability
- A measure of household composition
- Religion
- Born in the UK – Yes/No
At LA level the recommendation is for at least 5 people to share similar characteristics. This is not necessarily because of the ease of identification, but more because the information is extremely sensitive.
It is important to consider whether other information is available in the public domain that shows low counts of ethnic group and age. An example could be census tables at LA level. This could be used to pinpoint whether the respondent is likely to be the only one or one of a small number.
Choosing the ‘k’ value
It is very important to choose the ‘k’ value carefully. For example, 5 people sharing the same characteristics might not be enough if some LAs may have almost the full population surveyed and some of the demographic questions are very identifying. For example, the response “I live in a care home” could be very identifying and so could particular combinations of adults and children living in a household. The month and year of birth by LA, and detailed ethnicity could also lead to identification. This is despite the uncertainty from some 15-year-olds not responding.
Sample size
This is a survey with an unusually large sample: 40% of the population of 15-year-olds. Initial thoughts are that LA and month of birth are too detailed for a safeguarded dataset, particularly given the sensitivity of other variables.
Variables that should be omitted
We do not recommend including sexual identity in a safeguarded dataset. In addition, alleged criminal offences such as illegal drug use constitute sensitive personal data under the 2018 DPA. LA level is required but the combination of Index of Multiple Deprivation (IMD) quintile and LA can disclose a smaller geography of a Lower-layer Super Output Area (LSOA). This should only be included in a secure access version.
If a customer wants very detailed information on specific variables, it may be that they must commission this output separately to avoid putting all detailed variables in one file. Perturbation of some kind should be considered to avoid ‘differencing’ between outputs.
Record swapping
Record swapping by pairing similar 15-year-olds and swapping their records to each other’s LA does not solve all problems. It can create issues of perception of disclosure if the researcher is not aware that swapping has been carried out, so ensuring metadata and appropriate communication are very important here. Researchers may also perceive the data as damaged. The idea is to swap the types of individuals with the highest disclosure risk. Even if swapping is used, additional measures are recommended too.
Applying disclosure control
This is a highly sensitive dataset (especially the health and drug use questions) with identifying data at LA level. This means it should be released with greater protection than typically offered by a safeguarded file, which would require recoding and reduction of detail to reduce the risk of disclosure.
Initial ideas for applying disclosure control included:
- limiting the disability measure to two categories: ‘Has a disability? Yes, or No’
- recoding the household composition variable to give much less detail on relationship to other household members
- swapping LA codes of records to target people who are rare or unique within their area and are most at risk of re-identification – this would introduce doubt as to whether the records are allocated to the real geographic area
- ensuring the metadata provided with the dataset states that people with rare or unusual combinations of characteristics have been swapped between LAs
- swapping records with other records that that match on key variables – it is recommended matching variables are IMD quintile, sex and ethnic group as a minimum (ideally both LAs should be in the same Region)
Following further discussions with NHS Digital it was decided that:
- all potential identifying variables should be removed, apart from LA, gender, ethnic group (5 groups) and IMD (3 groups consisting of High, Medium, and Low)
- 25% of counts below 5 should be randomly selected to swap between LA, matching pairs to swap on gender, ethnic group and IMD – this is because counts below 5 are considered to be risky records on a combination of LA, gender, ethnic group (5 categories) and IMD (3 categories consisting of High, Medium, and Low)
Comments on the data
There was further discussion on the relationship between the microdata and small numbers in tables produced from the same data. The aggregated tables have been produced on the unswapped data. This would lead to some inconsistencies in outputs, but it was concluded that this would make it more difficult for a person to be identified.
There was a request to release the data as an open data set. If this were to happen, all Geography must be removed.
Sample size
This is the final size of the data after sampling and non-response are accounted for. Microdata based on a larger final sample will have a greater absolute risk of identification than a smaller sample. This is because the number of re-identifications is likely to be greater. Lower sample sizes should reduce the likelihood of identification risk if there is no response knowledge. If it is known that somebody has responded, the likelihood of identification will probably have increased depending on their characteristics. A smaller sample size is more helpful to an intruder when they know that a particular person is present in the data. This means that unless an intruder is likely to have response knowledge, microdata based on larger samples should be treated as having greater risk, and the riskiest scenario would be when a relatively large sample of a defined subgroup is taken. Microdata from these subsamples will be potentially more disclosive.
Suggestions of how to apply disclosure control to microdata from samples of different sizes are given later in this guidance.
Issues relating to household surveys
Many social surveys are household-based, such as the Transformed Labour Force Survey (LFS). Microdata from these surveys are hierarchical, as they include a record for each person in a household as well as variables which allow the records for each person to be linked into households. This enables an intruder to enhance identification keys, for example by combining information about the age, sex, marital status and relationship of each person in the household. Such keys increase the likelihood of households being identified since combinations of people are more likely to be unique than individual people on their own. This means the disclosure risk must be assessed at the individual and household level.
Large households
Households with many members (known as large households) increase the disclosure risk. Research by Duncan et al in 2011 demonstrates that the combination of single year of age and sex of all members of a household will be unique for most households above a relatively modest size, even without any detailed geography.
Longitudinal data
Some surveys use the same respondents for several successive periods. For instance, four waves of respondents may be used, with each wave contributing to four successive surveys and being replaced in turn over the four periods. Microdata from such surveys have increased disclosure risk because successive datasets may be combined to help identify a contributing household or person through an observed change in their circumstances.
For some longitudinal surveys there will be a requirement for safeguarded microdata to allow users to link sample members across successive years. Suggestions for dealing with longitudinal datasets are shown later in this guidance in the section titled “Dealing with large households”. This is a complex problem into which there is ongoing research and practical discussion:
- “Confidentiality challenges in releasing longitudinally linked data” – research by Lancaster University
- The Census 2021 Data Asset longitudinal data source for population in England and Wales: design and plans – a publication by the Office for National Statistics (ONS)
The Department for Digital, Culture, Media and Sport (DCMS) Taking Part Survey
The Taking Part Survey was a continuous face to face household survey of adults aged 16 and over, and children aged 5 to 15 years old in England. It ran from 2005 and was later replaced by the Participation Survey in 2021.
The survey collected data on participation in leisure, cultural and sporting activities in England, including:
- engagement in arts
- visits to heritage sites
- visits to museums and galleries
- use of public library services
The survey also collected a range of data covering:
- engagement with culture and sport whilst growing up
- volunteering
- charitable donations
- TV, radio and newspaper consumption
- internet use
- a range of socio-demographic variables
The data are longitudinal and include respondents who had been interviewed in at least one of the collections from 2011 to 2012 and following years.
Risk assessment and applying disclosure control
This risk assessment focuses on the data suitable for release under UK Data Service Safeguarded Licence (equivalent to End User Licence) and Special Licence (SL) releases. SL releases are an option for non-ONS datasets. These datasets are anonymised as with safeguarded data but contain more detail, which means they have a higher level of risk that might be reflected in any user agreement.
The inclusion of respondents in households with ten or more people is not a risk in this case as household size is top-coded at ‘3+ adults, 3+ children’. However, there are several factors that need to be considered in relation to the data, with associated actions to apply disclosure control.
Sampling
It is important to remember that primary sampling units (PSU) are mostly single postcode sectors. About 0.6% of households in a PSU (Primary Sampling Unit) will have been sampled. The sample size is small, and the method does allow the sample drawn from less densely populated PSUs to be reduced. This low sampling fraction means that only a small number of measures need to be applied for the Safeguarded release.
Longitudinal data
This is an unusual case of longitudinal data in that a person is tracked over several waves, rather than their household. Some waves will be missing for some people.
To protect against the disclosure of sensitive information from these risks, detailed linked information should be available only in the SL version, where access is more controlled and less readily obtained. Alternatively, it may be possible to use a safeguarded longitudinal dataset which provides less detail, and which cannot be linked on unique combinations of variable values to a more detailed cross-sectional dataset.
Geography
Region is generally the lowest geography used for safeguarded releases. Several area classification variables are proposed for inclusion along with region. Combinations of these variables with this relatively high-level geography may disclose a much lower level of geography, such as A Classification of Residential Neighbourhoods (ACORN).
For any combination of categories with fewer than 10 Output Areas in a Region, the codes will need to be suppressed until the resulting combination is not rare. If only one combination is rare, secondary suppression is required – or the removal of all geography. This should also be a consideration for SL as well, where the presence of LA will increase the likelihood of there being rare or unique combinations disclosing a lower geography.
Response knowledge
A further factor to consider is that of response knowledge. This occurs if a respondent can identify themselves, or if others in the household can identify a respondent from the data, knowing the respondent is present. It is difficult to quantify the risk, but unusual combinations of variables could lead to identification.
Socio-demographic variables
There are a few potential risky variables in this dataset, with different issues to consider and actions to take to apply disclosure control.
Age
The level of aggregation for this variable should be considered carefully. 5-year age bands top-coded at ‘80+’ are acceptable. Top-coding at 80 means that low counts associated with people over the age of 80 are avoided.
Marital status
Uniques may exist in specific categories. These need to be identified.
Former, separated and surviving civil partnership categories should be combined with their corresponding ‘married’ categories. This reduces the number of low counts.
Highest level of qualification
This variable needs to be carefully considered as certain categories may be too specific.
Keep six categories ranging from ‘degree level and above, or equivalent’ to ‘fewer than 5 GCSE A* to C, or equivalent’ for Safeguarded and SL versions. There are sufficient numbers in each category.
Health and disability
It is important to consider the level of detail which can be kept about specific conditions which are included in the data.
Any specific conditions or impairments should be included only in the SL dataset along with variables detailing reasons why activity is restricted, and these reasons include health problems. Safeguarded data should include only:
- general health – with responses such as ‘very good’, ‘good’, ‘fair’, ‘poor’, ‘very poor’
- an indicator of whether a person has a long-lasting health condition or illness
This is a very sensitive variable so limited detail included in the safeguarded file.
Employment
The survey collects very specific information on respondent’s and HRP employment, including write-ins of employer’s area of activity and their own job specific activity. This information is not proposed for inclusion in either the safeguarded or SL dataset and is used only to produce the socio-economic classification.
National Statistics Socio-Economic Classification (NS-SEC) operational and 8-category analytical classifications are acceptable. In the absence of any Standard Industrial Classification (SIC) variables, 4-digit SOC (Standard Occupational Classification) is also suitable for inclusion in the safeguarded release. As combinations of NS-SEC and SIC cannot be obtained, detailed NS-SEC can be included in the safeguarded data.
Life stage or life events
The life events coding frames provide a list of possible reasons why respondents may engage in a range of activities more often than they did previously, or less often. These include several reasons which can act to confirm tentative re-identifications and others which are sensitive.
You should only include these coding frames in the SL version of the data.
Free text responses
Free text responses are not suitable for inclusion in a safeguarded dataset. The respondent can state anything they wish in these replies, which could present a risk as they can contain information that would help with re-identification. They should be omitted or coded to broad categories.
Comments on the data
There are clear differences between the safeguarded and SL versions. The aim for SL data is to apply a lower level of disclosure control. The longitudinal relationships are weaker in the safeguarded data, which increases the level of confidentiality.
The disclosure risk has been reduced considerably by the disclosure applied to the socio-demographic variables. The relatively high level of Geography also reduces the risk for the safeguarded data.
DCMS had intended for the Taking Part survey to be conducted online in the future, if the survey had continued. This would have made it possible to follow up more respondents and more likely that respondents who move to live with a different household could be followed up. Vigilance would have been required due to a likely increase in the numbers who would be exposed to the risk of disclosure.
Applying statistical disclosure control
Once you have considered the likely intruder scenarios and identified the risk factors, the process of preparing microdata which are not personal information can be divided into three steps:
- Removing direct identifiers
- Applying disclosure control to key variables
- Dealing with large households
Removing direct identifiers
Direct identifiers are variables that allow a person to be identified from that information alone. These include variables such as:
- name
- address
- postcode
- National Insurance number
- NHS number
- Passport Number
If the data contain any direct identifiers, they must be removed. Information about dates of birth must also be removed, and it is generally advised that all similar variables (such as year of birth and month of birth) should also be removed. These will usually be substituted by age.
Removing direct identifiers does not necessarily provide sufficient protection. The data can still be personal information.
Unique and rare combinations of variables may still be present in the data, which will enable people or households with these characteristics to be identified.
Applying disclosure control to key variables
The data provider should consider:
- the relevant intruder scenarios
- what key variables are included in the data
- the size of the sample
The advice given in this section is based on experience gained from disclosure risk assessments which the Statistical Disclosure Control (SDC) Expert Group has carried out.
The NHS Digital Smoking Drink and Drugs Survey of teenagers
This has been released as a safeguarded dataset for years with no disclosure control necessary, other than basic anonymisation. The request is to see whether it can be released as an open data file (OGL).
The Smoking, Drink and Drugs Survey is a survey of 11 to 15-year-olds conducted at school in exam conditions. It includes questions on behaviour and attitudes concerning smoking, drinking and drugs. The “exam” is invigilated by an interviewer from a research company and the teachers are (ideally) kept out of the room. If they insist on being there, they are made to position themselves so they are unable to see what the children are writing.
The sample size is around 6,000 out of an eligible population of around 3 million. This equates to a 0.2% sample or 1 in 500, expanding to 17,500 so it will become a 0.6% sample (or 1 in 170). This case study is based on the 0.2% sample.
The only geographic code is Region. Variables that might significantly aid identification are:
- Region
- Gender
- School year (age)
- Ethnic Group (5 categories)
- Whether the child receives free school meals (FSM)
- How many people the child lives with
- A school ID which does not reveal the school but does connect a group of pupils from the same school
The sample contains unique combinations of region, gender, age, Ethnic Group and FSM. Sample uniques are not necessarily population uniques, although the position is more complex if a ‘motivated intruder’ knows a child has taken part in the survey.
Approximately 4% of records are unique and 15% have combination counts less than 5.
Similarly, if some of the behavioural information about a person is known, then it may be possible to identify them. For example, identification will be possible if a combination of taking a particular drug in a particular month is considered to be unique in terms of their region, gender, school year, ethnic group and FSM status.
If a parent knows their child took part in the survey, it is likely that they could identify them and therefore find out which drugs they are taking, or whether they smoke or drink alcohol.
There will be similar situations for other rare behaviours identified from this survey. Potential intruders could include fellow pupils, teachers and parents who are all likely to have a high response knowledge, meaning they will have a good idea of who has taken part in the survey. These people will know a lot of the identifying information about the children taking part and might be able to find behaviours which need to be protected.
There is a particular interest in the school pseudo ID. On its own it seems safe, but if one child is identified it makes it much easier to identify others at that school. It is useful for research as there is also a questionnaire which the schools complete on lesson provision around smoking, drinking and drugs. The school pseudo ID enables that information to be linked to the pupil level data. Some researchers also use it for multi-level modelling.
Risk assessment
As part of the assessment more detail was requested.
Selecting the sample
It is important to determine how children in the sample were selected. In particular, it is important to know how many schools were represented in the sample, and details of the range of the numbers sampled at each school.
In a typical year, 522 schools were randomly selected with 58 in each of the 9 England regions. There is sometimes slight oversampling in some regions to allow for differing response rates. The aim here is to have the same number of schools in each region in the achieved sample. School participation is low, with only 210 schools agreeing to take part. 35 children were then chosen within each school, with 7 children from each year group. Nearly all the selected children took part. This will change to choose three complete classes from different year groups (1 class from years 7 to 8 and two classes from years 9 to 11) within a school. This will mean around 90 children per school will take part. The overall achieved sample size is predicted to rise to around 17,500 per year.
When the survey look place
The research company conducted the ‘exams’ during the Autumn term.
Low cell counts
There were 942 records in cells representing fewer than 5 children. This was not from a cross tabulation of the 7 identification variables listed. The data excluded how many people the child lives with and the school ID. If the household size is included, then this increases to 2,663. If school ID is included then all except 50 records are in a cell size of less than 5. Note that the household size is not banded at all: 20% of respondents are in a household of 5 or more and 9% are in a household of 6 or more.
Other released tables
Data has been released as a National Statistics report with aggregate tables. The record level file was placed on the UK Data Service as a safeguarded file.
Applying disclosure control
For the open dataset, removing region is a necessity. Sensitive datasets, accessible to all, require stringent disclosure control to ensure an extremely low level of risk. This would probably be enough to deal with the risk of re-identification from behaviour described previously.
As mentioned, there is a risk of response knowledge where parents, teachers and other children know a particular child has taken part in the survey. There are additional risks arising from response knowledge, including:
- that a child who has completed the survey with a unique set of responses will identify themselves
- that a school with particularly pro-active approach to education on smoking, for example, could be identified from the information in the Teacher dataset (a school that lacks a pro-active approach to this kind of education could be identified in the same way)
- a school could be identified partly based on the number of children taking part in the survey – this could happen where this is lower than the selected sample
Any identification of a child immediately distinguishes the other records of pupils at the same school and this in turn increases their risk of re-identification.
For the open dataset the advice from the ONS was to:
- remove all geography
- include school year rather than age
- remove ‘…in last month’ variables relating to illegal drug use only and combining these with the ‘..in last 12 months’ variables
For the dataset released as a safeguarded dataset, the SDC team at the ONS noted that the variable list seemed to say that single year of age for most ages was provided, but the questionnaire did not reflect this. The advice from the SDC team was that information on the school year rather than the children’s age should be included.
Researchers were also advised that they should consider obtaining ethical approval for final release of the open data. This is because the data relate to children and contain a high level of detail on their use of alcohol and illegal drugs.
The presence of the school ID is a response knowledge issue. This risk is not normally considered for safeguarded datasets because of the licencing conditions. However greater caution is suggested in the case of datasets on children and for datasets which include sensitive information about people, such as the use of illegal drugs. It is even arguable that such information should be released under stricter controls. For this reason, it was suggested that a pseudo-identifier for schools is not included.
Comments on the data
NHS Digital decided to release the dataset under safeguarded conditions and not open access. They explained that the ONS had previously recommended that for a safeguarded file they should:
- Remove age but keep school year to mitigate against children being identified who were ahead or behind their natural school year
- Band all variables on household size to mitigate against very large households being identified
The report being published exclusively uses age rather than school year in the aggregated tables. Therefore, the preference is to do this in the safeguarded file and drop school year, rather than age. Note that everyone is coded as ages 11, 12, 13, 14 or 15, so any children who start secondary school early or leave late are recoded to 11 and 15 respectively.
This section lists key variables with suggestions of ways they may be protected for safeguarded data releases.
Note that these are suggestions and not rules, and this list is not exhaustive. The method of disclosure control you choose should be appropriate for the survey and the sample size. User requirements should always be considered. If a variable is needed at a lower level than advised, then another variable should be protected at a higher level. For example, if single year of age rather than banded age is required, then salary could be banded instead, or occupation and industry variables provided at a higher level. Having microdata supplied at a higher geography is always a good option for protecting confidentiality, but may not be viable in consideration of user needs. Where appropriate, each key variable is connected to the previously outlined scenarios of:
- use of public datasets
- spontaneous recognition
Specific variables could include the respondent’s residence or place of work.
A note on geography
Geography variables are the primary candidates for protection, as removing low levels of geography introduces extra uncertainty into a possible identification. For most safeguarded microdata, the lowest level of geography is Region. One of the main reasons for enabling access to Approved or Accredited Researchers under secure settings is to give researchers safe access to data with lower geographical details, such as Local Authority (LA) or Unitary Authority (UA). Some combinations of variables can imply a geography that is too low for a safeguarded dataset. An example is the urban/rural indicator which is based on postcode, so inclusion of this variable may reveal a lower level of geography. You should also be careful where other variables in the public domain may help with revealing a geography lower than Region, such as where microdata samples include Council Tax and associated variables. As LAs publish their rates of council tax, this could reveal the LA.
Where there are few areas of a particular category in a Region, geo-demographic segmentation classifications can help with identification, including:
- ACORN and MOSAIC
- the ONS Output Area Classification
- Index of Multiple Deprivation (IMD)
Records with such a category should have both the category and the Region suppressed. Suppressing only one might leave enough detail for an intruder to attempt to work out the missing value.
However, there will be some surveys for which it is appropriate to include LA or UA. An example would be where the primary purpose of the survey is to look at local matters requiring sub-regional data. In such cases, you should ensure the level of detail of other variables is correspondingly reduced to protect against identification.
Recoding highly identifiable variables (such as age of individual, or family or household structure) may enable data to be published at a more detailed level of geography. For example, ages could be banded into 10-year age groups, salaries could be banded, and information on variables such as family structure or number of children could be provided in coarser detail.
It is also important that variables which are based on postcode, such as urban or rural indicator and deprivation factor, are not included in such safeguarded datasets. Note that if variables such as deprivation factor are also essential for researchers, they can sometimes be represented by quintile or decile values. Each case needs to be considered on its own merit, but if LA or UA are included it is essential that no lower geography can be deduced from the data.
You can consult the Statistical Disclosure Control Expert Group for help or advice by emailing SDC.Queries@ons.gov.uk.
Reason for protection
These key variables must be protected because public datasets can be used to help identify people in the data.
Suggested treatment
The lowest geographical level will generally be English Region, (or Wales or Scotland as a country).
If data are required at a lower geography than these, it may be necessary to reduce the level of detail in some of the other variables.
Specific variables could include:
- respondent’s age
- age the respondent left full-time education
- age of the oldest child in a household under the age of 16
Reason for protection
These key variables must be protected because of previously mentioned intruder scenarios, the use of public datasets and spontaneous recognition.
Suggested treatment
For small samples, a single year of age can be provided.
For medium sized samples, ages should be banded – possibly into 5 year groups.
There are some rough guidelines around target sample sizes.
Small sample sizes
If the sample is less than 1% of the population or defined subpopulation, then most key variables may not need to be protected. However, there will always be some – such as geography – which need treatment.
Medium sample sizes
If the sample size is between 1% and 3% of the population, then it is likely that several key variables will need to be protected.
Large sample sizes
If the sample size is greater than 3% of the population, then further protection may be necessary. In general: the larger the sample size, the larger the possible percentage of a sub-population. In these cases, it may be necessary to remove some records or locally protect specific key variables. Local protection can be carried out by applying recoding or other techniques to particular categories, records or groups of cells only.
Suggested treatment
No additional protection is recommended for households of size lower than 10. You can find further information about this in the next section about “dealing with large households”.
Reason for protection
These variables must be protected because there are more than 250 possible values of these variables.
Suggested treatment
You should consider whether this level of detail is needed. It may be acceptable to band these variables to create categories such as UK, or EU.
Reason for protection
This variable can be categorised in several ways. The most basic categorisation would consist of 5 categories and 19 categories would be the most detailed set of standard categories in the harmonised question. There may also be write-in options.
Suggested treatment
You should decide on the level of detail which will be acceptable to users. If the user wants 19 categories, there will need to be additional reduction in the detail of other variables. Write-in options are not normally recommended for EUL unless there are large numbers of cases.
Specific variables may include:
- main job
- secondary job
- previous job
Reason for protection
These key variables must be protected because of the intruder scenarios, the use of public datasets and spontaneous recognition. Coding frames for these variables are generally to 3 digits or 4 digits.
If both occupation and industry are given to 4 digits, then the combination may be disclosive. For example, data about a company director of a particular manufacturer in a specific region. When Standard Industrial Classification (SIC) codes are revised, then if data have previously been published with low-level SIC codes, they should not be re-published with the new codes.
Suggested treatment
You should consider whether 4-digit level needs to be included. The general recommendation is that if industry is given to 4 digits, occupation should only be given to 3 digits, or the opposite way around.
Specific variables could include:
- gross and net salary
- annual, weekly, or hourly salary. bonuses
Reason for protection
Company directors’ salaries are in the public domain. Very high salaries and bonuses are often published.
Suggested treatment
Very high salaries and bonuses should be protected by top-coding at an appropriate level, for example at 10 times the average salary in the sample. Weekly and hourly rates will need to be correspondingly top-coded.
Specific variables could include:
- household income
- gross and net income
Reason for protection
This is a key variable in the intruder scenario around the use of public datasets.
Suggested treatment
These variables should at least be rounded to nearest £1,000. Very high values should be top-coded at an appropriate level, similarly to salaries. As an example, income variables for the Living Costs and Food Survey (LCFS) household level and person level datasets are top-coded to the value of the 96th percentile.
Reason for protection
These should all be considered. Examples include large lottery winnings, which may have been published in the media.
Council Tax rates are discoverable, so may help reveal a lower geography in combination with region or other geography variable.
Suggested treatment
Large lottery winnings should be top-coded at a defined value, such as £500,000 or £1,000,000.
Some social surveys include Council Tax band for the respondent along with the amount of Council Tax paid. There are also several variables based on the Council Tax payments.
Each Local Authority (LA) or Unitary Authority (UA) publish their rates of Council Tax at each band. This means the band plus the amount paid can lead to disclosure of a respondent’s location, either at the Authority level or lower, depending on the distribution of Council Tax rates across the authority. It is strongly recommended that safeguarded microdata should not include any geographical detail below region level. If both band and amount paid are included in the data, then there needs to be disclosure control of these.
There are different ways in which Council Tax payments and variables derived from them may be protected. The Living Costs and Food Survey (LCFS) team have developed a method that involves:
- Calculating an approximation by pooling several authorities within a region, from which a pool average is calculated.
- Applying these averages to each household so that the value of Council Tax shown is close to the real value, but that the local authority cannot be identified.
Sample sizes need to be considered alongside the likely response knowledge of the intruder. As previously stated, the larger the sample size, the greater the likelihood of an intruder with no response knowledge being able to identify a member of the dataset. However, a smaller sample would give less protection if an intruder knew a particular person was present in the sample.
Examples of protecting key variables
The list of key variables given in this guidance is not exhaustive and is provided only as a guide. Data providers are the experts on their data, which should mean they are able to confidently assess which variables may pose a risk. You should ensure that when a variable is protected all variables derived from it are similarly protected so that the original values cannot be discovered. You should also consider the number and detail of key variables in the dataset when applying protection.
Other key variables
Other variables such as tenure, number of children in household, marital status and qualifications are included in one or both of Scenarios 1 and 2. These need to be considered in the context of the survey and the likely requirements of users. They may additionally be candidates for protection when dealing with large households.
Article 9 of the UK-GDPR as enacted in the Data Protection Act defines the following variables as being “special categories of personal data”.
“Processing of personal data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, and the processing of genetic data, biometric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person’s sex life or sexual orientation shall be prohibited.”
If these variables are to be included in the microdata, you must consider the effect of any harm or distress that may be caused if an identification claim was to be made – whether that claim were to be correct or incorrect. For instance, a published claim that a named person was suffering from a particular illness or had been the victim of a particularly sensitive crime would be traumatic to the person concerned, whether true or not.
In addition, a variable such as Case Number could allow linkage with another dataset and you should examine any free text carefully to see what it reveals.
Civil partnerships and same-sex marriages have introduced possible new values to the marital status variables. Both are discoverable data and are therefore considered to be in the public domain.
As an example, inspection of a particular social survey dataset found that there were 2 respondents who stated that they were separated from their partner in a civil partnership, and 2 who had been in a civil partnership which had been legally dissolved. These small numbers made the microdata potentially disclosive. The ONS SDC Expert Group agreed that such values should be grouped together, to give a new value for the marital status variable of “civil partner or former civil partner”. This solution is recommended for all social surveys with variables including marital status or derived from it.
However, it is likely that, as the number of people in civil partnerships or same-sex marriages increases, the number of people in former relationships will also increase. This could mean that the risk of disclosure falls to an acceptable level. This means that each set of microdata should be considered on its own merits. A useful guide is that there should be at least 3 respondents in a Region for whom the variable has the same value.
In some cases, this could be regarded as the most sensitive of variables. It needs to be addressed as this question is included in some government surveys. Possible values of this variable are:
- heterosexual or straight
- gay or lesbian
- bisexual
- other sexual orientation
- “prefer not to say”
The disclosure risk here is that the numbers responding in some categories such as “bisexual” and “other” might be very small.
In theory, this would not be a problem unless data on sexual orientation has been published and could be taken together with safeguarded microdata to help identify someone. This is unlikely to happen. But it is also necessary to consider the risk of self-identification. For example, a researcher, knowing that their household was a respondent to the survey, might recognise themselves from the data by seeing information about sexual orientation and the combination of some other attributes. They could then discover the sexual orientation of another member of their own household. The SDC Expert Group does not normally advise protection against self-identification, as it would make it very difficult to release safeguarded datasets at all. In this case, however, the sensitivity of the variable makes it advisable to consider this risk. This is due to the likely perception of respondents that they could be identified, and their sexual orientation could be revealed, or that (if their sexual orientation is already widely known) sexual orientation could be used to help or confirm identification.
Variables that reveal the sexual orientation of respondents who are gay or lesbian but are not in a civil partnership or same-sex marriage need protection. In the Labour Force Survey (LFS), the variables on living together along with gender disclose same-sex couples who are not in civil partnerships. This means these should not be included in safeguarded datasets.
Recommendation
For safeguarded microdata, the sexual orientation variable should not be included.
It is now acceptable for bisexual to be included as a separate category in datasets in trusted research environments. Note that for those, there will be output checks to ensure outputs exported are of negligible risk.
A gender identity variable is included in Census 2021 data.
The risks of identification are likely to be similar to those for sexual orientation. As such, gender identity should not be included in any safeguarded release. However, as with sexual orientation, the variable can be included in both protected public tables and microdata released in safe settings to Approved or Accredited Researchers.
Further issues
Case numbers
The structure of case numbers (or serial numbers) in social survey microdata can reveal information about the geography of a household. Data providers should consider whether case numbers need to be included in safeguarded datasets. It is recommended that case numbers should be pseudonymised.
Free text
Free text can be described as personal details and opinions which can be included in the responses to non-specific questions in a survey. Free text variables should never be included untreated in safeguarded microdata. It may be feasible in some cases to publish free text variables in suitably anonymised form to Approved and Accredited Researchers for access in a secure setting. This could be achieved by categorising the responses, or by including key words only.
All proposed microdata releases – including training or teaching datasets – require a risk assessment.
The Office for National Statistics (ONS) Living Costs and Food Survey
The Living Costs and Food Survey (LCFS) collects information on spending patterns and the cost of living that reflects household budgets across the country. The primary uses of the survey are to provide information about spending patterns for the Consumer Price Indices, and about food consumption and nutrition.
Many user requests are made for microdata from this survey. This case study is about a specific request for ‘teaching microdata’, meaning microdata that would be made available ‘publicly’ under Open Government License (OGL). The OGL has few conditions other than not to misrepresent the data, as opposed to the End User License (EUL) where users must not try to identify an individual respondent, or to claim to have identified a respondent. A request for a teaching dataset therefore required considerable disclosure control, as this would be available without restriction.
Risk assessment
The following information was used in the risk assessment:
- the most likely intruder scenarios were considered to be ‘use of public datasets’ and ‘spontaneous recognition’ – both these scenarios are outlined in this guidance on microdata from surveys
- the sample size is about 0.04% of the UK population
- the survey is not longitudinal
- the survey is a household survey
Applying disclosure control
When applying disclosure control for this dataset it is important to consider that:
- this is a household dataset with individual-level variables for only the Household Reference Person
- the sample size is 0.04% households with the achieved sample being around 6,000 – the coverage is the UK
- few socio-demographic variables are included – sex, 3-class NS-SEC (National Statistics Socio-economic Classification), 4-category economic activity
- household size variables are top-coded at “4+ adults, 2+ children”
- gross household income is not banded, but is top-coded satisfactorily at approximately £62,000 per year – the main source is “earned” or “other”
- the only geographic variable is “modified” region, where London is split into Inner London and Outer London
The dataset included a case number, or record identifier. This should be independent for the open dataset and not relate to case numbers in any other version of the data.
Comments on the data
This is a training dataset for public use. This means only a relatively small amount of information is released. Only a small number of socio-demographic variables are included.
A relatively low value for top coded household income was used.
The household size is top-coded at “6”, which is normally suitable for a public use dataset.
The data will be of use to researchers and students to develop code which could then be run on a more detailed extract on a safeguarded dataset.
Dealing with large households
The risk of disclosure increases when each of the following applies:
- the survey is household based
- the microdata are hierarchical
- identification keys can be composed of variables such as the age and sex structure of the household and relationships of people to each other
If the size or composition of the household is also known by an intruder, this is likely to increase the disclosure risk and the probability of identification.
Extra protection might be given to some large households, as they usually contain both adults and children, and at the present time there are limited datasets which include detailed information for both adults and children.
Where large households consist of only adults (such as student-only households), the risk of matching with published data is likely to be mitigated by their increased mobility. This means that no additional protection is recommended for households with fewer than 10 residents.
It is possible that future development of commercially available datasets will increase the likelihood of being able to match them with records for large households with fewer than 10 residents. The SDC Expert Group will therefore keep this situation under review.
Longitudinal surveys
Longitudinal surveys, or surveys with a longitudinal element, such as the LFS and the Wealth and Assets Survey (WAS), pose a further risk. This is because linking successive waves or years can disclose more information about respondents than a cross-sectional survey. Microdata from such surveys may need to be subjected to further restrictions so that successive datasets cannot be linked, as this would increase disclosure risk to an unacceptable level. Possible methods are to have a higher level of banding on demographic variables such as marital status and number of children. For WAS, the geography variable was removed from the safeguarded dataset and the detail of ethnic group and religion was severely reduced. This ensured the variables that relate to wealth were retained in the data in detail, as these were of most interest.
If there is a requirement to allow safeguarded microdata to reflect the longitudinal nature of a survey by allowing individual households to be linked over time, then additional modification of the data is necessary. For example, a change in household size or marital status may allow an intruder to identify a household or individual by means of various pieces of information in the public domain, such as:
- births
- deaths
- marriages
- civil partnerships
- divorce registrations
In Wave 1 of a survey, a household has 3 residents. A is married to B and they have a child, C. Between Wave 1 and 2, A and B split up and D moves in. In Wave 2 of the survey, the household has 3 residents (A, D, and C).
If former resident B had access to the microdata – most likely through being an accredited researcher – they would be able to identify the household through knowledge of the household and find out information about individual D, or perhaps changes to A and C. Likewise, if D were an accredited researcher, they would be able to identify the household and find personal details about B from Wave 1 information.
Data providers should consider methods such as banding ages, marital status, socio-economic and relationship variables. Read more about confidentiality issues in longitudinal surveys.
In some circumstances users may require more details so they can study the change over time in respondents’ circumstances. In these cases, consideration should be given to supplying the data under the ONS Accredited Researcher protocol or allowing access through a secure environment in the relevant department.
Other approaches
You can consider modifying the database directly, especially where there are small numbers of records that have very unusual combinations of characteristics which may disproportionately affect the disclosure risk.
One approach is to remove the risky records, though you should be aware that this may make the microdata inconsistent with any other publications constructed from the microdata.
Alternatively, individual records that have unique or rare characteristics may be swapped with other records in other geographic areas. This has an advantage of retaining the usefulness of the data at higher geographies. Note that it is often beneficial to swap households rather than individual people to maintain relationships within households. This helps avoid the possibility of creating implausible inter-personal relationships. It is usual to match swapped records with others that have similar (but not identical) characteristics to ensure the data retains its usefulness.
Post tabular methods for tabular releases
It is increasingly common for users to create their own tabulations from the underlying microdata using some form of table builder. In many of these cases the microdata have been protected by the methods described in this guidance to ensure the resulting tables are not disclosive. For example, geography may have been coded to a higher level which reduces the risk of any variable published at the lowest geography helping to build the microdata record. Age can also be banded into groups. Read Stephanie Blanchard’s paper about the methodological challenges of protecting outputs from a flexible dissemination system, published in the Survey Methodology Bulletin (number 79).
An alternative approach is to apply a method which has been used to produce tables from previous Australian Censuses. This process adds changes or, ‘perturbations’, to the cell values following the assignment of a unique key value to each record. It introduces consistency so that a cell created from the same records will always have the same perturbation added. A similar approach is part of the methodology applied to Census 2021 in England and Wales. The SDC Expert Group can advise on the design and implementation for this approach if required.
A post-hoc method is intruder testing. You can find out more about this in the introduction to statistical disclosure control. Separate guidance on intruder testing is available on the ONS website. If an intruder can make lots of correct identifications, this would suggest disclosure control had been applied too weakly. However, if they are unable to make any correct identifications this could demonstrate that either the correct level of disclosure control has been applied, or that over-protective disclosure control has been applied.
It is possible to estimate the possibility of a sample unique being a population unique by determining the approximate number of people in the population with the characteristics of the sample member. From here, you can develop formal procedures for identifying the riskiness of a record and the dataset. Read more about privacy protection from sampling and perturbation in survey microdata.
However, intruder testing is very resource intensive. You should consider the balance of that level of resource and time against the possibility of using more straightforward disclosure protection.
Help and support
The aim when producing microdata to be released under licence is that the risk of disclosure is low and the data remain useful. Reducing risk completely is infeasible, as the resulting data would be so limited that they would be of little use. The defined level of risk is a subjective matter, and it is not straightforward to find a suitable level of disclosure control to apply. However, advice is available from the SDC Expert Group.
If you would like to discuss any issues related to applying statistical disclosure control to microdata from survey data, please contact the team by emailing SDC.Queries@ons.gov.uk.
Anybody requiring advice on releasing safeguarded level microdata can initiate the risk assessment process with the ONS SDC team by completing the social survey microdata checklist and sending it to the team by emailing SDC.Queries@ons.gov.uk. All government data should follow the advice in the DPA with reference to the variables defined as sensitive.
For queries on these case studies and submission of examples for possible addition to this document, please contact the SDC Expert Group by emailing SDC.Queries@ons.gov.uk.
See a flowchart of the SDC process from receiving data to output.