Confidentiality and Anonymity
The source treats these as two standards, not two words for the same promise. One controls who may see identifying information; the other removes your own ability to know who said what — and a single common design choice makes the second impossible.
Two standards, and the difference between them
The supplied teaching source reproduces an excerpt from a published research methods knowledge base which introduces this pair with a sentence worth reading literally: there are two standards that are applied in order to help protect the privacy of research participants. Two. Not one idea with two names.
The structural difference is where the identifying information sits. Under confidentiality it sits with you, protected by a control you have promised to operate. Under anonymity it was never created, so there is no control to operate and nothing to fail.
Confidentiality — access is controlled
- You know who each participant is
- Identifying information exists and is held somewhere
- The promise is about who it is released to
- The boundary is "directly involved in the study"
- The protection depends on your systems and your discipline holding for the life of the data
- It is compatible with follow-up, member checking, withdrawal of an individual's data, and repeated measurement
Anonymity — identity is never acquired
- You do not know who each participant is
- Identifying information was not collected, or was severed irreversibly
- The promise is about what exists at all
- The boundary includes you — the researcher is inside the exclusion
- The protection does not depend on anyone's later conduct
- It forecloses follow-up, member checking, individual withdrawal, and matching a person to themselves over time
Note the asymmetry in the second column. Anonymity is not simply a firmer version of confidentiality; it buys its strength by destroying a capability. Everything in that final row is something you will want at some point in a research project, and anonymity trades all of it away at once.
Why the stronger standard is often unavailable
The source does not merely say anonymity is hard. It names a specific structural cause: situations where participants have to be measured at multiple time points, with a pre-post study given as the example.
The logic is worth spelling out because it generalises. A pre-post design asks whether a change occurred within the same people. To answer that you must pair each person's later measurement with their own earlier one. Pairing requires a persistent handle on the individual — and a persistent handle is exactly what anonymity forbids.
This is not a limitation you can engineer around with care. It is a property of the question. Any design whose claim is about change within individuals, rather than change in a population aggregate, needs the ability to match a participant to themselves, and so cannot be anonymous. The design families this affects are set out in Correlation and Experimental Methodologies.
WHERE ANONYMITY SURVIVES AND WHERE IT DOES NOT
| What your design needs to do | Anonymity available? | Why |
|---|---|---|
| Measure a population once and report aggregates | Yes | Nothing has to be linked to anything. A single unlinked response set is the natural home of the stronger standard |
| Measure the same people before and after an intervention | No | The source's own named exception. Pairing each person's two measurements requires a handle that persists across both |
| Interview participants and return transcripts for checking | No | You cannot return a transcript to a person you cannot identify |
| Allow an individual to withdraw their data after collection | No | Withdrawal requires locating that person's records among the rest |
| Follow up non-respondents or send a reminder | No | Knowing who has not responded means knowing who is who |
| Combine a survey with later interviews of selected respondents | No | Selection and recontact both need identity to persist beyond the first instrument |
| Analyse existing organisational records with names stripped before you receive them | Yes, sometimes | Only if the severing happens before the data reaches you and cannot be reversed by anyone who will hand you more |
The pre-post row is the source's stated example. The remaining rows apply the same reasoning to other designs and are synthesis added here.
The three identifiability classes an application must declare
The abstract choice between two standards becomes concrete at one line of the ethics application, and it is that line rather than the vocabulary that will govern your project.
The three classes
- Individually identifiable
- The data carries information that identifies a specific person. Names, employee numbers, role titles in a small team, or any combination that resolves to one individual.
- Re-identifiable (coded)
- Identifiers have been replaced with a code, and a separate key exists which allows the code to be resolved back to the person. Identity is removed from the dataset but is not destroyed.
- Non-identifiable
- The data has never been identified, or identifiers have been removed such that no specific individual can be identified. There is no key.
The middle class is the operational compromise, and it is where most workplace project research belongs. Coding gives you a dataset that is safe to analyse, safe to show a supervisor, and safe to quote from, while the key preserves the capability that pure anonymity destroys — the ability to pair, follow up, and honour a withdrawal request.
MAPPING THE STANDARD ONTO THE CLASS
| Privacy standard you are offering | Identifiability class you declare | What must be true for the declaration to hold |
|---|---|---|
| Confidentiality, with identities held throughout | Individually identifiable | Access is restricted to those directly involved in the study, and you can name them |
| Confidentiality, with identities separated from the data | Re-identifiable (coded) | The key is stored apart from the coded data, with its own access list, and it is retained or destroyed on a stated schedule |
| Anonymity | Non-identifiable | No key exists anywhere, including in your own records, your recruitment list, or the platform that collected the responses |
The three classes are the source's; the mapping between standard and class is analysis added here to make the choice operable.
Confidentiality is three separate promises, not one
The application does not ask whether your study is confidential. It asks you to describe how you will preserve participants' confidentiality as you collect and analyse the data and when you report the results. Three moments, named separately, with different failure modes.
The three moments and what each one asks
While the data is being gathered
Who else is present or can overhear. Where the interview happens and whether being seen entering the room is itself disclosure. What the recruitment mechanism reveals — a calendar invitation, a room booking, a distribution list.
While you are working with it
Who has access to raw transcripts and recordings as against coded extracts. Whether the working files carry names. Whether a supervisor, a translator, a transcription service or a co-marker sits inside or outside "directly involved in the study".
When the results are written and read
Whether a quotation identifies its speaker to a colleague. Whether a table with small cells resolves to one person. Whether the combination of role, tenure and project named across three sentences describes exactly one individual.
For as long as the data is retained
Not one of the three the source names, but implied by the retention requirement. The promise has to survive a minimum five-year storage period, staff changes, and any later access you granted in the application.
The reporting moment is the one most often underestimated, because it is the only one where the audience includes people who already know the participants. Anyone can de-identify a transcript against a stranger. The test that matters is whether a person's own manager, reading the finished report, can tell who said it.
Choosing your standard and writing it down
Which standard your design can actually support
Whichever standard you land on, it has to be stated in the same words in three places: the application, the participant information sheet, and the consent record. Divergence between them is the version of this failure that examiners and committees actually catch. The full structure those statements sit inside is in Writing an Ethics Application, and the information-sheet obligations are in Voluntary Participation and Informed Consent.
Evidence that your privacy standard is real
- One named standard — confidentiality or anonymity — used consistently in the application, the information sheet and the consent record
- One declared identifiability class, chosen before collection begins
- Where data is coded: the key stored separately from the data, with its own access list and its own disposal date
- A written list of everyone who will see raw data, including transcribers, translators, supervisors and second coders
- A stated rule for how quotations will be attributed in the report
- A check that no reported table, split or subgroup resolves to a single individual
- A storage and retention plan that carries the promise through the full retention period
What to carry forward
- Confidentiality controls who may see identifying information. Anonymity means it never existed, including for you. The source treats them as two standards, and so should your application.
- Anonymity is stronger because it removes the possibility of failure rather than managing it — and it pays for that by destroying the ability to link, follow up, check back and withdraw.
- The source names the structural reason anonymity is often impossible: participants measured at multiple time points. Any within-person claim rules it out.
- Re-identifiable (coded) data is the working compromise for most project research: identity out of the dataset, key held separately, capability preserved.
- Declare the identifiability class before collection. It can be lowered later and never raised.
- Confidentiality is required separately at collection, analysis and reporting. Reporting is where small populations give people away through combinations rather than names.
Frequently asked questions
Can I promise anonymity in a workplace survey?
Only if the instrument genuinely cannot resolve a response to a person — no response identifiers, no gated links, no invitation list you can cross-reference, and no follow-up round. If you need to remind non-respondents, pair two waves, or let someone withdraw their answers, you are offering confidentiality with coded data and should say so in those words.
What is the difference between coded data and anonymous data?
A key. Re-identifiable (coded) data has had identifiers replaced with a code while a separate key still allows the code to be resolved back to the individual. Non-identifiable data has no key anywhere. The distinction is not about how the dataset looks but about whether the link to the person still exists somewhere in your project.
Does my supervisor count as someone 'directly involved in the study'?
The source defines the confidentiality boundary as anyone not directly involved in the study, and the application separately asks you to specify who apart from yourself and your supervisors will have access to the data and results. That phrasing treats supervisors as inside the boundary by default. Anyone else — a transcription service, a translator, a second coder, a workplace sponsor — needs naming and justifying.
Why does a pre-post design make anonymity impossible?
Because the design's claim is about change within the same individuals, which requires pairing each person's later measurement with their own earlier one. Pairing needs a handle that persists across both rounds, and a persistent handle is exactly what anonymity excludes. The supplied source names this case specifically as the reason anonymity is sometimes difficult to accomplish.
How do I stop a quotation identifying its speaker?
The supplied source states the requirement to preserve confidentiality when reporting results but supplies no technique for it. In practice the effective controls are aggregating role detail, removing project-specific and idiomatic phrasing, and testing every quotation against the question of whether the speaker's own team leader could name them. Where a passage cannot be protected, seek specific consent for that passage or leave it out.
Which identifiability class should I declare if I am not sure?
Work backwards from the capabilities your design needs rather than from how private it feels. If you need to link, recontact or honour a withdrawal, you need a key, and the honest declaration is re-identifiable. Declaring non-identifiable to look rigorous and then discovering you needed the key is not recoverable, because the class can be lowered later but never raised.
References and source attribution
- Trochim, W. M. K. 2006, Research Methods Knowledge Base (http://www.socialresearchmethods.net/kb/probform.php) — the source of the excerpt on the system of ethical protections, including both privacy standards quoted on this page.
- Cooper, D. & Schindler, P. 2008, Business Research Methods, 10th ed., McGraw-Hill — the text the supplied source draws on for its treatment of research ethics.
- Quinlan, C. 2011, Business Research Methods, 1st ed., Cengage Publishing, Chapter 3 — the prescribed reading accompanying the ethics material.
- Australian Code for the Responsible Conduct of Research (2007) — named in the supplied source as the instrument behind the identifiability classification and the data retention requirements. The code itself is not reproduced in the supplied material.
- The supplied teaching source: consolidated weekly teaching notes and slide material on research ethics, including the eight-part ethics application structure reproduced as an example from a government department.
Suggested questions for Ask KEVOS
- My study measures the same team before and after a process change. Which identifiability class should I declare, and where should the key live?
- Check whether my survey can honestly be described as anonymous given how I am distributing it.
- Draft the confidentiality section of my application covering collection, analysis and reporting as three separate statements.
- My participants are eight people on one delivery team. How do I report subgroup results without resolving to individuals?
- List everyone who will touch my raw data and tell me which of them need naming in the access section.
