Suggestion: improve AI processing pipeline, parallelism and hardware utilization
Lc31 Dns
lc31dns at gmail.com
Mon Sep 14 17:36:43 BST 2026
Hello,
I would like to suggest a performance improvement for digiKam’s AI
analysis, especially Object Detection / Auto-Tags when processing very
large collections recursively.
I am currently using digiKam 9.2 on Kali Linux with an Intel Core i7-4770,
16 GB of RAM and an NVIDIA GeForce RTX 3060 12 GB.
My use case is demanding: I run the analysis recursively over an entire
disk containing a very large and heterogeneous collection of files, rather
than processing only a few individual folders.
I have performed extensive tests to verify the hardware and GPU
acceleration.
The RTX 3060 is working correctly. digiKam correctly loads OpenCV DNN,
OpenCL and the NVIDIA/CUDA libraries, opens the NVIDIA devices, and the
digiKam process itself really performs compute work on the RTX 3060.
Using nvidia-smi
pmon during actual Object Detection processing, I measured digiKam reaching
approximately 70–86% GPU SM utilization at times.
However, utilization is very irregular. GPU load repeatedly drops
significantly or becomes idle while the analysis is still running. At the
same time, CPU, system RAM and GPU resources are not continuously used near
their available capacity.
The issue therefore appears to be mainly related to the processing pipeline
and insufficient parallelism between its different stages, rather than a
lack of GPU support.
Operations such as file reading, image decoding, preprocessing, GPU
inference, post-processing and database writing should overlap much more
efficiently instead of one stage frequently waiting for another.
A more efficient architecture could use:
parallel file reading -> parallel decoding/preprocessing -> queued/batched
GPU inference -> parallel post-processing -> asynchronous/batched database
writes
While the GPU is processing one batch, CPU threads should already prepare
the following batches, and previous results should be written to the
database concurrently. RAM could be used for prefetching and buffering
prepared work, and GPU batches/queues could be adjusted so that the GPU
remains supplied with data whenever possible.
The amount of parallelism should scale according to the hardware available:
CPU cores, RAM, GPU performance, VRAM and storage performance.
In addition, the user should be able to control how aggressively digiKam
uses the computer. A simple global performance setting would be enough, for
example:
Low / Normal / Maximum
In Maximum mode, the user would explicitly authorize digiKam to use as much
CPU, RAM, GPU, VRAM and I/O capacity as reasonably possible, with maximum
practical parallelism. This would be especially useful for unattended jobs,
for example when starting a complete recursive analysis overnight and
wanting digiKam to use the full available performance of the machine until
the task is finished.
On a smaller computer, the user could choose a lower level and keep more
resources available for other applications.
The important point is not simply to “use more GPU”. The goal is to improve
the processing pipeline, increase parallelism between stages, keep work
queued ahead of the GPU, scale according to the available hardware, and
allow the user to decide how much of the machine digiKam may use.
The current GPU acceleration works, but the overall processing architecture
does not appear to scale efficiently enough with powerful hardware during
very large recursive AI analysis jobs.
Thank you for considering this improvement.
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://mail.kde.org/pipermail/digikam-users/attachments/20260914/a3629ea2/attachment.htm>
More information about the Digikam-users
mailing list