A person seated at a desk with multiple monitors displaying code and maps, working in a dimly lit office, blurred face, and a mug nearby.

Automating Face and Plate Blurring via REST API and Docker: Batch Pipelines for High-Volume Teams

Mateusz Zimoch
Published: 9/5/2026
Updated: 9/24/2026

TL;DR: Once redaction volume passes a few hundred files a month, the bottleneck stops being detection speed and becomes queue design, retries and quality gates. Gallio PRO offers a Docker container, a Linux command line interface and a REST API on the Enterprise plan, so anonymization can sit inside your own pipeline. This is for MLOps engineers, platform teams and dataset owners building that pipeline.

Three signs the desktop application has stopped scaling

Every anonymization workflow starts the same way: someone opens a file, processes it, checks it, exports it. That works, until three symptoms appear together.

The queue is a person. Work arrives faster than one operator clears it, and prioritisation lives in somebody's inbox rather than in a system.

There is no retry. A job that fails at 80 percent of a four hour file is restarted by hand, from the beginning, by whoever notices.

There is no record. Nobody can answer which files were processed last quarter, with which settings, and who approved the export. That is an accountability problem before it is an engineering one, because GDPR Article 5(2) requires the controller to be able to demonstrate compliance, not merely to achieve it.

People working in a control room, using multiple screens and electronic equipment, focused on monitoring various digital displays.

Two integration patterns

Pattern A: watched folder plus container

Files land in an ingest directory. A container picks them up, processes them, writes output to a separate directory and moves the source to a holding area. Scheduling is handled by cron, systemd or your orchestrator.

This suits teams with predictable batch windows, for example a nightly run against the day's camera exports. It is the lowest effort pattern and needs no application development. Gallio PRO supports it through the Docker container and the Linux command line interface available on the Enterprise plan.

Pattern B: REST API called by your system

Your case management system, evidence platform or data labelling tool submits a job and receives the processed result back. Redaction becomes a step in an existing workflow instead of a separate tool a human remembers to use.

This suits teams where redaction is triggered by an event, such as an export request being approved, and where the requesting system already owns identity, permissions and audit. Request the API documentation from the vendor before you design against it, since endpoint names, authentication method, payload limits and whether processing is synchronous or job based all shape the client you write.

Person working on a laptop in an industrial setting, examining diagrams, surrounded by machinery and equipment.

The pipeline in seven stages

Stage

What it does

What it protects against

1. Ingest

Accept file, record checksum and source

Silent duplicates and unexplained files

2. Validate

Check format, resolution, duration, codec

Jobs that fail two hours in

3. Queue

Persist job state with priority

Work lost on restart

4. Process

Run detection and blurring

The part everyone already thinks about

5. QA gate

Sample and review before release

Missed detections reaching a recipient

6. Export

Deliver to the requester, log the release

Untracked disclosure

7. Purge

Delete the original on schedule

Article 5(1)(e) storage limitation breaches

Stages 2, 5 and 7 are the ones teams skip and later rebuild. Supported input formats are JPG and PNG for images and MP4 and MOV for video, so validation at stage 2 is a cheap check that saves expensive failures.

Five engineering details that decide reliability

  1. Idempotency by content hash. Key jobs on the checksum of the source file, not the filename. Camera exports collide on names constantly.
  2. Segment long files. Split multi hour recordings into chunks, process, then reassemble. A failure costs you one chunk rather than the whole job.
  3. Bound concurrency to licences. Floating licensing on the Enterprise plan determines how many workers can run at once. Set the worker pool to that number rather than to your core count.
  4. Separate scratch from delivery. Partial output must never be writable to the directory a requester collects from. This is how unredacted or half redacted material escapes.
  5. Log jobs, not people. Record file hash, settings, timestamps and operator identity. Do not build a log of what was detected. Gallio PRO itself stores no detection logs and no personal data, and your pipeline should not reintroduce that risk.
Black and white image of server racks with code overlay, giving a digital and technological atmosphere.

The QA gate is not optional

Automation changes who does the checking, not whether checking happens. Build an explicit gate with three rules: a fixed sampling rate for routine batches, a full review for anything going to a court, a regulator or the public, and a defined route into the manual editor for exceptions.

That exception route matters because automatic detection covers faces and license plates only. Name badges, monitor screens, documents, tattoos and logos are redacted in the built in manual editor, so any file where those appear leaves the automated path by design. Flag those categories at ingest if you can, using camera location or request type, so they are routed rather than discovered at the gate.

For volume specific guidance, see large scale anonymization for big sets of photos and videos. If the footage originates in a video management system, the integration considerations are covered in video anonymization in VMS systems.

How to design the pipeline in six steps

  1. Write the state machine before the code. List every job state and every transition, including failure and manual review. Most pipeline defects are missing states, not bad algorithms.
  2. Request the API documentation and the container specification. Confirm authentication, payload limits, synchronous versus job based processing, and exit codes before committing to a design.
  3. Benchmark one worker on your own footage. Use the free Gallio PRO demo, which is unlimited and watermarked at up to 720p, to establish throughput for a single worker, then multiply by your licensed concurrency.
  4. Define the QA gate and its sampling rate. Write down who reviews, at what rate, and what happens when a miss is found. Include the manual editor route for badges, screens and documents.
  5. Instrument the pipeline without personal data. Track throughput, failure rate, queue depth and time to delivery using file identifiers, never detection contents.
  6. Automate the purge. Schedule deletion of originals and document the schedule. An automated pipeline that never deletes anything is a growing liability, not an efficiency gain.
Black and white image of a computer screen displaying system performance data and numerical statistics, with a dark background.

Where this fits, and where it does not

This architecture suits recorded material processed in batches: camera exports, evidence packages, archive backlogs and training datasets. For image heavy dataset work, the same patterns apply through Gallio PRO image anonymization, typically with much higher file counts and shorter per file processing.

It does not suit live feeds. Gallio PRO processes recorded files and does not perform live stream processing, so a pipeline design that assumes a continuous stream needs a different tool category. Say that early in the design review rather than discovering it during integration.

Because everything runs on your own infrastructure, footage never leaves your environment and no external processing dependency enters your architecture, which also keeps the compliance surface small. Deployment options are on the Gallio PRO video anonymization page.

Close-up of a row of black server racks in a data center, with dim lighting highlighting metal components.

FAQ

Does Gallio PRO have an API for automating anonymization?

Yes. A REST API, a Docker container and a Linux command line interface are available on the Enterprise plan. Request the API documentation from the vendor before designing your client.

Can anonymization run in a container on our own servers?

Yes. The Docker container is the usual choice for a shared internal processing node, with processing happening entirely on your own hardware.

How do I handle video files that are hours long?

Segment them into chunks, process each chunk independently and reassemble. A failure then costs one chunk rather than the entire job, and progress survives restarts.

Do I still need human review if the pipeline is automated?

Yes. Build an explicit QA gate with a sampling rate for routine batches, full review for material going to courts, regulators or the public, and a manual editor route for objects that are not detected automatically.

Can anonymization be applied to a live camera stream?

Not with Gallio PRO. It processes recorded files rather than live streams, so pipelines should be designed around batch or event triggered processing of stored material.

Mateusz Zimoch

CEO and Co-Founder of Gallio PRO, is an engineer with extensive expertise in Computer Science, Data Science, Robotics, and Artificial Intelligence. Mateusz graduated with dual majors from the Wrocław University of Technology and launched two startups centered on developing AI-powered Computer Vision solutions and constructing Remote Operated Vehicles (ROVs).