KEVOS
ArticlesServicesCase studiesAboutContact
ArticlesServicesCase studiesAboutContact
← ArticlesConfidentiality and AnonymityProject Delivery · Research ProjectsLesson 91/115← PrevNext →
GuidePublished 16 Aug 202615 min readBy KEVOS Editorialconfidentiality and anonymity researchanonymity stronger standard privacycoded research datapre-post study anonymity
On this page

Ask about this page

KEVOS AIConfidentiality and Anonymity

KEVOS knowledge first · trusted web sources when needed

KEVOS/Project Delivery/Research Projects/Research Ethics
Project DeliveryResearch ProjectsCoreResearch Ethics

Confidentiality and Anonymity

The source treats these as two standards, not two words for the same promise. One controls who may see identifying information; the other removes your own ability to know who said what — and a single common design choice makes the second impossible.

Reading time17 minutes
LevelCore
Topic streamResearch Ethics
Source materialResearch Ethics
Updated2026-08-16

In brief

  • The source names two standards applied to protect participant privacy. They are not synonyms and they are not interchangeable in an application.
  • Confidentiality means identifying information will not be made available to anyone not directly involved in the study. The information exists; access to it is controlled.
  • Anonymity means the participant remains anonymous throughout the study — even to the researchers themselves. The information does not exist to be leaked.
  • The source calls anonymity the stronger guarantee and says plainly that it is sometimes difficult to accomplish, naming the reason: participants measured at multiple time points.
  • The operational form of the choice is a three-way classification your application must declare — individually identifiable, re-identifiable (coded), or non-identifiable.

Two standards, and the difference between them

The supplied teaching source reproduces an excerpt from a published research methods knowledge base which introduces this pair with a sentence worth reading literally: there are two standards that are applied in order to help protect the privacy of research participants. Two. Not one idea with two names.

From the source

Both standards, as stated

Confidentiality: "Almost all research guarantees the participant's confidentiality — they are assured that identifying information will not be made available to anyone who is not directly involved in the study."

Anonymity: "The stricter standard is the principle of anonymity which essentially means that the participant will remain anonymous throughout the study — even to the researchers themselves. Clearly, the anonymity standard is a stronger guarantee of privacy, but it is sometimes difficult to accomplish, especially in situations where participants have to be measured at multiple time points (e.g., a pre-post study)."

The structural difference is where the identifying information sits. Under confidentiality it sits with you, protected by a control you have promised to operate. Under anonymity it was never created, so there is no control to operate and nothing to fail.

Confidentiality — access is controlled

  • You know who each participant is
  • Identifying information exists and is held somewhere
  • The promise is about who it is released to
  • The boundary is "directly involved in the study"
  • The protection depends on your systems and your discipline holding for the life of the data
  • It is compatible with follow-up, member checking, withdrawal of an individual's data, and repeated measurement

Anonymity — identity is never acquired

  • You do not know who each participant is
  • Identifying information was not collected, or was severed irreversibly
  • The promise is about what exists at all
  • The boundary includes you — the researcher is inside the exclusion
  • The protection does not depend on anyone's later conduct
  • It forecloses follow-up, member checking, individual withdrawal, and matching a person to themselves over time

Note the asymmetry in the second column. Anonymity is not simply a firmer version of confidentiality; it buys its strength by destroying a capability. Everything in that final row is something you will want at some point in a research project, and anonymity trades all of it away at once.

Note

Confidentiality is also a named misconduct

"Breaking participation confidentiality" is the second item on the list of unethical behaviours the source supplies. It is one of only two items on that list that concern participants at all.

That matters for how you read the standard. Confidentiality is not a courtesy extended to participants; a breach is categorised alongside deception and misrepresentation. See Unethical Behaviour in Research.

Why the stronger standard is often unavailable

The source does not merely say anonymity is hard. It names a specific structural cause: situations where participants have to be measured at multiple time points, with a pre-post study given as the example.

The logic is worth spelling out because it generalises. A pre-post design asks whether a change occurred within the same people. To answer that you must pair each person's later measurement with their own earlier one. Pairing requires a persistent handle on the individual — and a persistent handle is exactly what anonymity forbids.

This is not a limitation you can engineer around with care. It is a property of the question. Any design whose claim is about change within individuals, rather than change in a population aggregate, needs the ability to match a participant to themselves, and so cannot be anonymous. The design families this affects are set out in Correlation and Experimental Methodologies.

WHERE ANONYMITY SURVIVES AND WHERE IT DOES NOT

What your design needs to doAnonymity available?Why
Measure a population once and report aggregatesYesNothing has to be linked to anything. A single unlinked response set is the natural home of the stronger standard
Measure the same people before and after an interventionNoThe source's own named exception. Pairing each person's two measurements requires a handle that persists across both
Interview participants and return transcripts for checkingNoYou cannot return a transcript to a person you cannot identify
Allow an individual to withdraw their data after collectionNoWithdrawal requires locating that person's records among the rest
Follow up non-respondents or send a reminderNoKnowing who has not responded means knowing who is who
Combine a survey with later interviews of selected respondentsNoSelection and recontact both need identity to persist beyond the first instrument
Analyse existing organisational records with names stripped before you receive themYes, sometimesOnly if the severing happens before the data reaches you and cannot be reversed by anyone who will hand you more

The pre-post row is the source's stated example. The remaining rows apply the same reasoning to other designs and are synthesis added here.

Caution

The word "anonymous" on an instrument you cannot make anonymous

The most common privacy failure in workplace research is not a breach. It is a survey headed "anonymous" that carries a response identifier, a login-gated link, an email invitation traceable to a list, or a free-text box in which people describe their own role.

Each of those reintroduces identity. If you offer anonymity and the instrument does not deliver it, you have made a promise the design cannot keep — and you have made it in writing, to people who relied on it.

The three identifiability classes an application must declare

The abstract choice between two standards becomes concrete at one line of the ethics application, and it is that line rather than the vocabulary that will govern your project.

From the source

The declaration the application requires

"Describe how, where and in what form the data will be stored and whether the data will be individually identifiable; re-identifiable (coded) data; or non-identifiable data."

The supplied source reproduces this as part of an eight-part application structure, introduced as an example sourced from a government department. The identifiability classification and the retention periods that sit beside it are presented as requirements flowing from a named national code for the responsible conduct of research — not as that form's local convention.

The three classes

Individually identifiable
The data carries information that identifies a specific person. Names, employee numbers, role titles in a small team, or any combination that resolves to one individual.
Re-identifiable (coded)
Identifiers have been replaced with a code, and a separate key exists which allows the code to be resolved back to the person. Identity is removed from the dataset but is not destroyed.
Non-identifiable
The data has never been identified, or identifiers have been removed such that no specific individual can be identified. There is no key.

The middle class is the operational compromise, and it is where most workplace project research belongs. Coding gives you a dataset that is safe to analyse, safe to show a supervisor, and safe to quote from, while the key preserves the capability that pure anonymity destroys — the ability to pair, follow up, and honour a withdrawal request.

MAPPING THE STANDARD ONTO THE CLASS

Privacy standard you are offeringIdentifiability class you declareWhat must be true for the declaration to hold
Confidentiality, with identities held throughoutIndividually identifiableAccess is restricted to those directly involved in the study, and you can name them
Confidentiality, with identities separated from the dataRe-identifiable (coded)The key is stored apart from the coded data, with its own access list, and it is retained or destroyed on a stated schedule
AnonymityNon-identifiableNo key exists anywhere, including in your own records, your recruitment list, or the platform that collected the responses

The three classes are the source's; the mapping between standard and class is analysis added here to make the choice operable.

Practice note

Decide the class before collection, not after

The source does not prescribe this; it follows from what the classes are. Identifiability can be lowered after collection and never raised. Data gathered as non-identifiable cannot later be linked to a follow-up round, and a key destroyed cannot be reconstructed.

So the class is a design decision made at the same moment as the sampling and the instrument, and written into the application before anyone is approached. Treating it as an administrative field to be completed at the end of the form is how projects lose a second data round they had planned all along.

Confidentiality is three separate promises, not one

The application does not ask whether your study is confidential. It asks you to describe how you will preserve participants' confidentiality as you collect and analyse the data and when you report the results. Three moments, named separately, with different failure modes.

The three moments and what each one asks

COLLECTION

While the data is being gathered

Who else is present or can overhear. Where the interview happens and whether being seen entering the room is itself disclosure. What the recruitment mechanism reveals — a calendar invitation, a room booking, a distribution list.

ANALYSIS

While you are working with it

Who has access to raw transcripts and recordings as against coded extracts. Whether the working files carry names. Whether a supervisor, a translator, a transcription service or a co-marker sits inside or outside "directly involved in the study".

REPORTING

When the results are written and read

Whether a quotation identifies its speaker to a colleague. Whether a table with small cells resolves to one person. Whether the combination of role, tenure and project named across three sentences describes exactly one individual.

AFTERWARDS

For as long as the data is retained

Not one of the three the source names, but implied by the retention requirement. The promise has to survive a minimum five-year storage period, staff changes, and any later access you granted in the application.

The reporting moment is the one most often underestimated, because it is the only one where the audience includes people who already know the participants. Anyone can de-identify a transcript against a stranger. The test that matters is whether a person's own manager, reading the finished report, can tell who said it.

Caution

Deductive disclosure in a small organisation

Identity leaks through combinations, not through names. "The scheduler on the eastern package" is a name if there is one of them. A table showing responses split by function and seniority can resolve to a single cell of one. A quotation containing an idiom, a technical preference or a well-known grievance identifies its speaker to everyone on the team.

Project research runs on small, tightly connected populations, which makes this the ordinary case rather than the edge case. The supplied source names confidentiality as a requirement and does not supply a de-identification technique or a minimum cell size; that judgement is yours to make and to defend.

Choosing your standard and writing it down

Which standard your design can actually support

IfYou need to compare each participant with themselves at two or more time points
ThenAnonymity is unavailable. Declare re-identifiable (coded) data, and describe where the key lives and who holds it.
IfYou will return transcripts to participants for checking, or allow individual withdrawal after collection
ThenAnonymity is unavailable for the same reason. Coded is the honest declaration; say so on the information sheet rather than implying more.
IfYou are running a single unlinked instrument, with no follow-up and no recontact
ThenAnonymity is available. Then check the instrument actually delivers it — no response identifiers, no gated links, no traceable invitations.
IfYour participants are colleagues and the population is small
ThenWhatever class you declare, plan reporting-stage protection separately. The class governs your files; it does not govern what a reader can deduce from your findings.
IfA transcription service, translator or second coder will see raw data
ThenThey are either directly involved in the study or they are a disclosure. Name them in the access section of the application and bind them to the same terms.
IfYou are unsure whether a variable is identifying
ThenAsk whether it can be combined with anything else you are publishing. Identity is a property of combinations, so test the variable against the whole reported set, not on its own.

Whichever standard you land on, it has to be stated in the same words in three places: the application, the participant information sheet, and the consent record. Divergence between them is the version of this failure that examiners and committees actually catch. The full structure those statements sit inside is in Writing an Ethics Application, and the information-sheet obligations are in Voluntary Participation and Informed Consent.

Evidence that your privacy standard is real

  • One named standard — confidentiality or anonymity — used consistently in the application, the information sheet and the consent record
  • One declared identifiability class, chosen before collection begins
  • Where data is coded: the key stored separately from the data, with its own access list and its own disposal date
  • A written list of everyone who will see raw data, including transcribers, translators, supervisors and second coders
  • A stated rule for how quotations will be attributed in the report
  • A check that no reported table, split or subgroup resolves to a single individual
  • A storage and retention plan that carries the promise through the full retention period
Check before you proceed

The colleague test

Take your most quotable finding and write it as it would appear in the report. Now hand it, mentally, to the participant's own team leader.

Can they name the speaker? If yes, confidentiality has failed at the reporting stage no matter how carefully the file was coded. Rewrite the attribution, aggregate the detail, or go back to the participant for specific consent to be identifiable in that passage.

Source gap

What the source states and what it leaves to you

The two standards are defined and the three identifiability classes are named as a required declaration. Beyond that, the supplied material is silent on method. There is no de-identification procedure, no minimum reportable cell size, no guidance on pseudonyms or on how quotations should be attributed, and no account of what to do if a participant is recognised after publication.

The national code named as the source of the identifiability requirement is not itself reproduced in the supplied material, so the definitions above are the source's summary of it rather than the code's own text.

The operational detail — storage medium, access control, key management, destruction — is likewise absent, and is discussed for what it is worth in Research Data Storage, Retention and Ownership. Get the technique from the body that will assess your application; take the standards and the classification from here.

What to carry forward

  1. Confidentiality controls who may see identifying information. Anonymity means it never existed, including for you. The source treats them as two standards, and so should your application.
  2. Anonymity is stronger because it removes the possibility of failure rather than managing it — and it pays for that by destroying the ability to link, follow up, check back and withdraw.
  3. The source names the structural reason anonymity is often impossible: participants measured at multiple time points. Any within-person claim rules it out.
  4. Re-identifiable (coded) data is the working compromise for most project research: identity out of the dataset, key held separately, capability preserved.
  5. Declare the identifiability class before collection. It can be lowered later and never raised.
  6. Confidentiality is required separately at collection, analysis and reporting. Reporting is where small populations give people away through combinations rather than names.

Frequently asked questions

Can I promise anonymity in a workplace survey?

Only if the instrument genuinely cannot resolve a response to a person — no response identifiers, no gated links, no invitation list you can cross-reference, and no follow-up round. If you need to remind non-respondents, pair two waves, or let someone withdraw their answers, you are offering confidentiality with coded data and should say so in those words.

What is the difference between coded data and anonymous data?

A key. Re-identifiable (coded) data has had identifiers replaced with a code while a separate key still allows the code to be resolved back to the individual. Non-identifiable data has no key anywhere. The distinction is not about how the dataset looks but about whether the link to the person still exists somewhere in your project.

Does my supervisor count as someone 'directly involved in the study'?

The source defines the confidentiality boundary as anyone not directly involved in the study, and the application separately asks you to specify who apart from yourself and your supervisors will have access to the data and results. That phrasing treats supervisors as inside the boundary by default. Anyone else — a transcription service, a translator, a second coder, a workplace sponsor — needs naming and justifying.

Why does a pre-post design make anonymity impossible?

Because the design's claim is about change within the same individuals, which requires pairing each person's later measurement with their own earlier one. Pairing needs a handle that persists across both rounds, and a persistent handle is exactly what anonymity excludes. The supplied source names this case specifically as the reason anonymity is sometimes difficult to accomplish.

How do I stop a quotation identifying its speaker?

The supplied source states the requirement to preserve confidentiality when reporting results but supplies no technique for it. In practice the effective controls are aggregating role detail, removing project-specific and idiomatic phrasing, and testing every quotation against the question of whether the speaker's own team leader could name them. Where a passage cannot be protected, seek specific consent for that passage or leave it out.

Which identifiability class should I declare if I am not sure?

Work backwards from the capabilities your design needs rather than from how private it feels. If you need to link, recontact or honour a withdrawal, you need a key, and the honest declaration is re-identifiable. Declaring non-identifiable to look rigorous and then discovering you needed the key is not recoverable, because the class can be lowered later but never raised.

References and source attribution

  1. Trochim, W. M. K. 2006, Research Methods Knowledge Base (http://www.socialresearchmethods.net/kb/probform.php) — the source of the excerpt on the system of ethical protections, including both privacy standards quoted on this page.
  2. Cooper, D. & Schindler, P. 2008, Business Research Methods, 10th ed., McGraw-Hill — the text the supplied source draws on for its treatment of research ethics.
  3. Quinlan, C. 2011, Business Research Methods, 1st ed., Cengage Publishing, Chapter 3 — the prescribed reading accompanying the ethics material.
  4. Australian Code for the Responsible Conduct of Research (2007) — named in the supplied source as the instrument behind the identifiability classification and the data retention requirements. The code itself is not reproduced in the supplied material.
  5. The supplied teaching source: consolidated weekly teaching notes and slide material on research ethics, including the eight-part ethics application structure reproduced as an example from a government department.

Suggested questions for Ask KEVOS

  • My study measures the same team before and after a process change. Which identifiability class should I declare, and where should the key live?
  • Check whether my survey can honestly be described as anonymous given how I am distributing it.
  • Draft the confidentiality section of my application covering collection, analysis and reporting as three separate statements.
  • My participants are eight people on one delivery team. How do I report subgroup results without resolving to individuals?
  • List everyone who will touch my raw data and tell me which of them need naming in the access section.

Related KEVOS knowledge

Voluntary Participation and Informed ConsentCore · research ethicsResearch Data Storage, Retention and OwnershipCore · research ethicsProtecting Participants from HarmCore · research ethicsWriting an Ethics ApplicationCore · research ethicsDependent Groups and Unequal PowerAdvanced · research ethicsThe Four Ethical PrinciplesCore · research ethics
KEVOS® · Project Delivery · Research Projects Page KVS-PM-RES-0091 · v1.0.0 · content 2026.08 Last reviewed 2026-08-16

Continue learning

Voluntary Participation and Informed ConsentGuide · Research ProjectsNEXT LESSON →Protecting Participants from HarmGuide · Research ProjectsThe Four Ethical PrinciplesGuide · Research ProjectsEthics Review and Approval ProcessesGuide · Research Projects
KEVOS · Engineering, manufacturing and project improvement
ArticlesServicesCase studiesAboutContact
© 2026 KEVOS®