ctctexttensordec

Tensor decoder element for CTC text recognition models. This element performs CTC decoding on inference output tensors and attaches the decoded text to the buffer as analytics metadata.

An example PP-OCRv5 recognition model and dictionary can be found at: https://huggingface.co/monkt/paddleocr-onnx

 gst-launch-1.0 filesrc location=/TEXT/IMAGE.png \
     ! pngdec ! videoconvert ! videoscale ! video/x-raw,pixel-aspect-ratio=1/1 \
     ! onnxinference model-file=/PATH/TO/languages/english/rec.onnx \
     ! ctctexttensordec dictionary-file=/PATH/TO/languages/english/dict.txt implicit-space=true \
     ! fakesink

This takes a PNG, performs text recognition on it via onnxinference, and decodes the inferred tensors with ctctexttensordec.

Hierarchy

GObject
    ╰──GInitiallyUnowned
        ╰──GstObject
            ╰──GstElement
                ╰──GstBaseTransform
                    ╰──ctctexttensordec

Factory details

Authors: – Seungha Yang

Classification:Tensordecoder/Video

Rank – primary

Plugin – rsanalytics

Package – gst-plugin-analytics

Pad Templates

sink

video/x-raw(ANY):
         format: { AYUV_F32, RGBA_F32LE, ARGB_F32, RGBA_F32BE, A444_16LE, A444_16BE, Y416_LE, AYUV64, RGBA64_LE, ARGB64, ARGB64_LE, BGRA64_LE, ABGR64_LE, Y416_BE, RGBA64_BE, ARGB64_BE, BGRA64_BE, ABGR64_BE, A422_16LE, A422_16BE, A420_16LE, A420_16BE, A444_12LE, GBRA_12LE, A444_12BE, GBRA_12BE, Y412_LE, Y412_BE, A422_12LE, A422_12BE, A420_12LE, A420_12BE, RGBA_F16LE, RGBA_F16BE, A444_10LE, GBRA_10LE, A444_10BE, GBRA_10BE, A422_10LE, A422_10BE, A420_10LE, A420_10BE, BGR10A2_LE, RGB10A2_LE, Y410, A444, GBRA, AYUV, VUYA, RGBA, RBGA, ARGB, BGRA, ABGR, A422, A420, AV12, RGBP_F32LE, RGBP_F32BE, RGB_F32LE, RGB_F32BE, Y444_16LE, GBR_16LE, Y444_16BE, GBR_16BE, Y216_LE, Y216_BE, v216, P016_LE, P016_BE, Y444_12LE, GBR_12LE, Y444_12BE, GBR_12BE, I422_12LE, I422_12BE, Y212_LE, Y212_BE, I420_12LE, I420_12BE, P012_LE, P012_BE, RGBP_F16LE, RGBP_F16BE, RGB_F16LE, RGB_F16BE, Y444_10LE, GBR_10LE, Y444_10BE, GBR_10BE, BGR10x2_LE, RGB10x2_LE, r210, I422_10LE, I422_10BE, NV16_10LE40, NV16_10LE32, Y210, UYVP, v210, I420_10LE, I420_10BE, P010_10LE, NV12_10LE40, NV12_10LE32, P010_10BE, MT2110R, MT2110T, NV12_10BE_8L128, NV12_10LE40_4L4, Y444, BGRP, GBR, RGBP, NV24, v308, IYU2, RGBx, xRGB, BGRx, xBGR, RGB, BGR, Y42B, NV16, NV61, YUY2, YVYU, UYVY, VYUY, I420, YV12, NV12, NV21, NV12_16L32S, NV12_32L32, NV12_4L4, NV12_64Z32, NV12_8L128, Y41B, IYU1, YUV9, YVU9, BGR16, RGB16, BGR15, RGB15, RGB8P, GRAY_F32LE, GRAY_F32BE, GRAY16_LE, GRAY16_BE, GRAY_F16LE, GRAY_F16BE, GRAY10_LE16, GRAY10_LE32, GRAY8 }
          width: [ 1, 2147483647 ]
         height: [ 1, 2147483647 ]
      framerate: [ 0/1, 2147483647/1 ]
        tensors: "tensorgroups\,\ ctc-text-recognition-out\=\(/uniquelist\)\{\ \(caps\)\"tensor/strided\\\,\\\ tensor-id\\\=\\\(string\\\)\\\{\\\ ctc-text-recognition-out-prob\\\,\\\ ctc-text-recognition-out-logits\\\ \\\}\\\,\\\ dims\\\=\\\(int\\\)\\\<\\\ \\\[\\\ 0\\\,\\\ 2147483647\\\ \\\]\\\,\\\ \\\[\\\ 0\\\,\\\ 2147483647\\\ \\\]\\\,\\\ \\\[\\\ 1\\\,\\\ 2147483647\\\ \\\]\\\ \\\>\\\,\\\ dims-order\\\=\\\(string\\\)row-major\\\,\\\ type\\\=\\\(string\\\)float32\"\ \}\;"

Presencealways

Directionsink

Object typeGstPad


src

video/x-raw(ANY):
         format: { AYUV_F32, RGBA_F32LE, ARGB_F32, RGBA_F32BE, A444_16LE, A444_16BE, Y416_LE, AYUV64, RGBA64_LE, ARGB64, ARGB64_LE, BGRA64_LE, ABGR64_LE, Y416_BE, RGBA64_BE, ARGB64_BE, BGRA64_BE, ABGR64_BE, A422_16LE, A422_16BE, A420_16LE, A420_16BE, A444_12LE, GBRA_12LE, A444_12BE, GBRA_12BE, Y412_LE, Y412_BE, A422_12LE, A422_12BE, A420_12LE, A420_12BE, RGBA_F16LE, RGBA_F16BE, A444_10LE, GBRA_10LE, A444_10BE, GBRA_10BE, A422_10LE, A422_10BE, A420_10LE, A420_10BE, BGR10A2_LE, RGB10A2_LE, Y410, A444, GBRA, AYUV, VUYA, RGBA, RBGA, ARGB, BGRA, ABGR, A422, A420, AV12, RGBP_F32LE, RGBP_F32BE, RGB_F32LE, RGB_F32BE, Y444_16LE, GBR_16LE, Y444_16BE, GBR_16BE, Y216_LE, Y216_BE, v216, P016_LE, P016_BE, Y444_12LE, GBR_12LE, Y444_12BE, GBR_12BE, I422_12LE, I422_12BE, Y212_LE, Y212_BE, I420_12LE, I420_12BE, P012_LE, P012_BE, RGBP_F16LE, RGBP_F16BE, RGB_F16LE, RGB_F16BE, Y444_10LE, GBR_10LE, Y444_10BE, GBR_10BE, BGR10x2_LE, RGB10x2_LE, r210, I422_10LE, I422_10BE, NV16_10LE40, NV16_10LE32, Y210, UYVP, v210, I420_10LE, I420_10BE, P010_10LE, NV12_10LE40, NV12_10LE32, P010_10BE, MT2110R, MT2110T, NV12_10BE_8L128, NV12_10LE40_4L4, Y444, BGRP, GBR, RGBP, NV24, v308, IYU2, RGBx, xRGB, BGRx, xBGR, RGB, BGR, Y42B, NV16, NV61, YUY2, YVYU, UYVY, VYUY, I420, YV12, NV12, NV21, NV12_16L32S, NV12_32L32, NV12_4L4, NV12_64Z32, NV12_8L128, Y41B, IYU1, YUV9, YVU9, BGR16, RGB16, BGR15, RGB15, RGB8P, GRAY_F32LE, GRAY_F32BE, GRAY16_LE, GRAY16_BE, GRAY_F16LE, GRAY_F16BE, GRAY10_LE16, GRAY10_LE32, GRAY8 }
          width: [ 1, 2147483647 ]
         height: [ 1, 2147483647 ]
      framerate: [ 0/1, 2147483647/1 ]

Presencealways

Directionsrc

Object typeGstPad


Properties

beam-width

“beam-width” guint

Number of prefixes to keep during beam search

Flags : Read / Write

Default value : 10


blank-index

“blank-index” guint

Index of the CTC blank token

Flags : Read / Write

Default value : 0


decoding-method

“decoding-method” Ctc-text-tensor-dec-method *

CTC decoding method

Flags : Read / Write

Default value : beam-search (0)


dictionary-file

“dictionary-file” gchararray

Dictionary file with one UTF-8 token per line

Flags : Read / Write

Default value : NULL


implicit-space

“implicit-space” gboolean

Append an implicit space token after the dictionary entries

Flags : Read / Write

Default value : false


top-k

“top-k” guint

Number of token candidates to consider at each beam search step (0 = all)

Flags : Read / Write

Default value : 0


Named constants

Ctc-text-tensor-dec-method

Members

beam-search (0) – BeamSearch: Decode using CTC prefix beam search
greedy (1) – Greedy: Decode by selecting the most probable token at each timestep

The results of the search are