Video anonymization in scientific research – definition
Video anonymization in scientific research is a set of technical and organizational measures used to prepare research recordings and images in a way that limits or eliminates the possibility of identifying individuals visible in the material. In practice, this mainly concerns faces, license plates, and other elements that reveal identity when they appear in the image and have identifying significance in a given research context.
In the context of visual materials, it is important to distinguish anonymization from pseudonymization. Under the GDPR, pseudonymization still constitutes the processing of personal data because the data can be linked to a specific person using additional information. Anonymization, by contrast, is irreversible. In research practice, full video anonymization is often difficult because identification may result not only from a face, but also from the scene context, voice, clothing, location, time of the event, or file metadata. For this reason, many datasets remain materials with a reduced risk of identification even after faces and license plates have been blurred, rather than data that can always be considered definitively anonymous.
The legal basis for this assessment is Regulation (EU) 2016/679 of the European Parliament and of the Council, namely the GDPR, applicable since 25 May 2018, as well as guidelines from the European Data Protection Board concerning, among other things, pseudonymization, data minimization, and privacy by design. In scientific research, this is supplemented by ethics approvals and ethics committee opinions, research institution policies, funder requirements, and standards for archiving and sharing research data.
The role of video anonymization in scientific research
In scientific research, video recordings and photographs are used in fields including the social sciences, medicine, psychology, transport, ergonomics, education, sports science, and behavioral research. Such material can be highly valuable for analysis, while at the same time containing personal data or special category personal data if it shows health status, disability, children, or intimate situations.
Video anonymization serves several purposes at once:
- implementing the data minimization principle under Article 5(1)(c) of the GDPR,
- reducing the risk of infringing the rights and freedoms of research participants,
- enabling safer sharing of materials with the research team, auditors, or data archives,
- meeting the conditions of ethics approval where a committee requires de-identification before analysis or publication,
- preparing material for secondary research use in line with institutional policy.
In image-based research, simply removing file metadata is not enough. If a face, vehicle registration number, or another unique identifier is visible in the recording, the material may still enable identification. That is why techniques that directly modify the image layer are essential.
How research photo and video anonymization works
The process should be planned before data collection begins. In practice, this means describing the sources of identification risk, establishing the legal basis for processing, and determining which image elements will be subject to automatic masking and which will require manual masking.
A typical process includes the following stages:
- Inventory of the material – photos, video sequences, file formats, metadata, recording length, and number of people.
- Risk classification – whether the material shows children, patients, individuals in sensitive situations, or private locations.
- Detection of identifying objects – primarily faces and license plates.
- Masking – most commonly through blurring, pixelation, or permanent occlusion of the area.
- Quality control – verifying that no frames remain with a visible face or plate.
- Archiving – separating the source version from the anonymized version and assigning access rights and retention periods.
In Gallio PRO, automatic blurring applies to faces and license plates. The software does not automatically detect logos, tattoos, personal ID badges, documents, or content displayed on monitors. These elements can be blurred manually in the built-in editor. This is important in field, laboratory, and observational research, where the source of identification may extend beyond a person’s likeness.
Technologies used for video anonymization
Effective anonymization of visual material requires accurate object detection in every frame or in representative image frames. In practice, machine learning models are used, most often from the deep learning domain. This is necessary because classical methods based solely on image features are usually less robust to changes in lighting, camera angle, partial face occlusion, or motion.
From a technical perspective, the process usually works as follows:
- a detection model locates a face or license plate in the image,
- a tracking algorithm links detections across consecutive frames,
- a masking module applies a blur or occlusion effect to the entire detected area,
- an operator reviews the material and makes manual corrections.
CNN architectures and their newer variants are used to build detection models, and object detection tasks also rely on one-stage and two-stage models. What matters most, however, is not which architecture was used, but whether the model achieves sufficient sensitivity and stability on data similar to those used in the study. Material from fixed cameras, smartphones, body-worn cameras, and laboratories has different characteristics, so the model should be validated for the specific use case.
An important practical limitation is that Gallio PRO does not perform real-time anonymization or live video stream anonymization. It is a solution for processing recorded material, which in scientific research is usually consistent with the actual research workflow and archiving procedure.
Key parameters and metrics for video anonymization
The quality of video anonymization should not be assessed solely based on a declaration that the material has been blurred. Measurable detection and quality control indicators are needed. In practice, both computer vision metrics and operational indicators related to dataset processing are used.
Parameter
Meaning
Why it matters in research
Recall
The proportion of actual faces or license plates detected by the system
Low recall means a risk of leaving visible personal data in the material
Precision
The proportion of correct detections among all detections
Low precision increases the number of false masks and the time needed for manual correction
IoU
Intersection over Union for the detected bounding area
It affects whether the mask covers the entire face or plate area
Miss rate per 1,000 frames
The number of undetected objects per 1,000 frames
Helps assess residual risk in long recordings
Processing time
Minutes of material processed per hour of system operation
Affects project planning and operating costs
Manual correction rate
The share of recordings requiring operator intervention
Shows process maturity and the actual workload involved
When assessing risk, a simple operational formula can be used:
Residual risk = P(non-detection) x Impact of disclosure x Exposure of the material
This is not a normative formula, but it is a useful working model for a DPIA and for research project assessment. Each component should be described qualitatively or quantitatively in line with the institution’s methodology.
GDPR, ethics approvals, and archiving research materials
In scientific research, image anonymization alone does not remove legal obligations at earlier stages of the data lifecycle. If the source material contains personal data, its collection, storage, and processing must have a legal basis. In practice, this will most often be a task carried out in the public interest, participant consent, or another basis provided for under national law and the institution’s internal regulations.
The most important requirements include:
- Article 5 GDPR – processing principles, including lawfulness, data minimization, purpose limitation, integrity, and confidentiality,
- Article 9 GDPR – special categories of data, where a recording reveals, for example, health status,
- Article 25 GDPR – privacy by design and privacy by default,
- Article 32 GDPR – appropriate security measures,
- Article 35 GDPR – a data protection impact assessment where the risk is high.
In scientific projects, ethics approvals and ethics committee opinions are also important. Ethics documentation should specify whether the material will be archived, for how long, who will have access to the source version, and whether the data will be shared with secondary researchers. A good practice is to store the raw version in a highly secure environment and provide only the anonymized version, or at least a version with reduced identification risk, for team analysis and publication.
In the archiving context, it is worth referring to the FAIR principles for research data, published in 2016, with the reservation that accessibility and interoperability must not undermine privacy protection. In practice, this means separating the descriptive layer from files containing visual material and controlling repository access rights.
Challenges and limitations of anonymizing visual materials
The biggest challenge is that a person may be identified from a combination of features. Blurring the face is not always sufficient. In a study, unusual clothing, location, voice, the known time of an event, or the presence of other people may all enable identification. Therefore, deciding whether material is truly anonymous requires contextual assessment.
There are also important legal and interpretive limitations. In many European countries, blurring license plates is generally treated as a precautionary standard and reflects national regulatory practice and a broad approach to privacy protection. In Poland, the status of license plates is less clear-cut. On the one hand, the Polish data protection authority and parts of EU case law point to the need for cautious treatment of indirect identifiers. On the other hand, Polish Supreme Administrative Court case law includes the view that a license plate does not always constitute personal data in itself. In scientific research, the safer approach is to follow the precautionary principle and mask license plates, especially when publishing material.
A person’s likeness also requires separate attention. Obligations arise not only under the GDPR, but also under the Civil Code and the Copyright and Related Rights Act. In practice, there are exceptions relating to public figures, wider crowd scenes, and situations where a person has been paid to pose, but their application always requires case-by-case analysis of the research purpose and the intended further use of the material.
Examples of video anonymization in scientific research
Photo and video anonymization is particularly useful wherever the image is needed to analyze a phenomenon, but the person’s identity is not needed by the researcher or should be restricted to a narrow group of authorized individuals.
- Pedestrian and driver behavior studies – masking faces and license plates before traffic analysis.
- Classroom interaction analysis – blurring students’ faces before providing material to external coders.
- Medical and rehabilitation research – limiting patient identification in recordings of exercises or consultations.
- UX and eye-tracking research in the lab – protecting participant identity while preserving behavioral and movement data.
- Field research in public spaces – publishing illustrative materials only after masking people and vehicles.
Normative references and sources
The following acts and documents provide the basis for interpreting the concept and practice of anonymizing visual materials in scientific research:
- Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 – GDPR.
- Article 29 Working Party, Opinion 05/2014 on Anonymisation Techniques, 10 April 2014.
- European Data Protection Board – GDPR guidance, privacy by design guidance, consent guidance, and guidance on the processing of personal data for scientific research purposes, where applicable.
- OECD Recommendation of the Council concerning Access to Research Data from Public Funding, 2021.
- Wilkinson et al., The FAIR Guiding Principles for scientific data management and stewardship, Scientific Data 3, 2016.
- Act of 4 February 1994 on Copyright and Related Rights.
- Civil Code – provisions concerning personal rights, including image rights and privacy.
If an institution applies local policies for research data management, repositories, and retention, these should be treated as a complementary layer alongside generally applicable law. In practice, these policies often define minimum archiving periods, the access model, and documentation requirements for video materials in greater detail.
See also
- Data anonymization
- Face blurring
- License plate blurring
- GDPR compliance