ctctexttensordec
Tensor decoder element for CTC text recognition models. This element performs CTC decoding on inference output tensors and attaches the decoded text to the buffer as analytics metadata.
An example PP-OCRv5 recognition model and dictionary can be found at: https://huggingface.co/monkt/paddleocr-onnx
gst-launch-1.0 filesrc location=/TEXT/IMAGE.png \
! pngdec ! videoconvert ! videoscale ! video/x-raw,pixel-aspect-ratio=1/1 \
! onnxinference model-file=/PATH/TO/languages/english/rec.onnx \
! ctctexttensordec dictionary-file=/PATH/TO/languages/english/dict.txt implicit-space=true \
! fakesink
This takes a PNG, performs text recognition on it via onnxinference,
and decodes the inferred tensors with ctctexttensordec.
Hierarchy
GObject ╰──GInitiallyUnowned ╰──GstObject ╰──GstElement ╰──GstBaseTransform ╰──ctctexttensordec
Factory details
Authors: – Seungha Yang
Classification: – Tensordecoder/Video
Rank – primary
Plugin – rsanalytics
Package – gst-plugin-analytics
Pad Templates
sink
video/x-raw(ANY):
format: { AYUV_F32, RGBA_F32LE, ARGB_F32, RGBA_F32BE, A444_16LE, A444_16BE, Y416_LE, AYUV64, RGBA64_LE, ARGB64, ARGB64_LE, BGRA64_LE, ABGR64_LE, Y416_BE, RGBA64_BE, ARGB64_BE, BGRA64_BE, ABGR64_BE, A422_16LE, A422_16BE, A420_16LE, A420_16BE, A444_12LE, GBRA_12LE, A444_12BE, GBRA_12BE, Y412_LE, Y412_BE, A422_12LE, A422_12BE, A420_12LE, A420_12BE, RGBA_F16LE, RGBA_F16BE, A444_10LE, GBRA_10LE, A444_10BE, GBRA_10BE, A422_10LE, A422_10BE, A420_10LE, A420_10BE, BGR10A2_LE, RGB10A2_LE, Y410, A444, GBRA, AYUV, VUYA, RGBA, RBGA, ARGB, BGRA, ABGR, A422, A420, AV12, RGBP_F32LE, RGBP_F32BE, RGB_F32LE, RGB_F32BE, Y444_16LE, GBR_16LE, Y444_16BE, GBR_16BE, Y216_LE, Y216_BE, v216, P016_LE, P016_BE, Y444_12LE, GBR_12LE, Y444_12BE, GBR_12BE, I422_12LE, I422_12BE, Y212_LE, Y212_BE, I420_12LE, I420_12BE, P012_LE, P012_BE, RGBP_F16LE, RGBP_F16BE, RGB_F16LE, RGB_F16BE, Y444_10LE, GBR_10LE, Y444_10BE, GBR_10BE, BGR10x2_LE, RGB10x2_LE, r210, I422_10LE, I422_10BE, NV16_10LE40, NV16_10LE32, Y210, UYVP, v210, I420_10LE, I420_10BE, P010_10LE, NV12_10LE40, NV12_10LE32, P010_10BE, MT2110R, MT2110T, NV12_10BE_8L128, NV12_10LE40_4L4, Y444, BGRP, GBR, RGBP, NV24, v308, IYU2, RGBx, xRGB, BGRx, xBGR, RGB, BGR, Y42B, NV16, NV61, YUY2, YVYU, UYVY, VYUY, I420, YV12, NV12, NV21, NV12_16L32S, NV12_32L32, NV12_4L4, NV12_64Z32, NV12_8L128, Y41B, IYU1, YUV9, YVU9, BGR16, RGB16, BGR15, RGB15, RGB8P, GRAY_F32LE, GRAY_F32BE, GRAY16_LE, GRAY16_BE, GRAY_F16LE, GRAY_F16BE, GRAY10_LE16, GRAY10_LE32, GRAY8 }
width: [ 1, 2147483647 ]
height: [ 1, 2147483647 ]
framerate: [ 0/1, 2147483647/1 ]
tensors: "tensorgroups\,\ ctc-text-recognition-out\=\(/uniquelist\)\{\ \(caps\)\"tensor/strided\\\,\\\ tensor-id\\\=\\\(string\\\)\\\{\\\ ctc-text-recognition-out-prob\\\,\\\ ctc-text-recognition-out-logits\\\ \\\}\\\,\\\ dims\\\=\\\(int\\\)\\\<\\\ \\\[\\\ 0\\\,\\\ 2147483647\\\ \\\]\\\,\\\ \\\[\\\ 0\\\,\\\ 2147483647\\\ \\\]\\\,\\\ \\\[\\\ 1\\\,\\\ 2147483647\\\ \\\]\\\ \\\>\\\,\\\ dims-order\\\=\\\(string\\\)row-major\\\,\\\ type\\\=\\\(string\\\)float32\"\ \}\;"
src
video/x-raw(ANY):
format: { AYUV_F32, RGBA_F32LE, ARGB_F32, RGBA_F32BE, A444_16LE, A444_16BE, Y416_LE, AYUV64, RGBA64_LE, ARGB64, ARGB64_LE, BGRA64_LE, ABGR64_LE, Y416_BE, RGBA64_BE, ARGB64_BE, BGRA64_BE, ABGR64_BE, A422_16LE, A422_16BE, A420_16LE, A420_16BE, A444_12LE, GBRA_12LE, A444_12BE, GBRA_12BE, Y412_LE, Y412_BE, A422_12LE, A422_12BE, A420_12LE, A420_12BE, RGBA_F16LE, RGBA_F16BE, A444_10LE, GBRA_10LE, A444_10BE, GBRA_10BE, A422_10LE, A422_10BE, A420_10LE, A420_10BE, BGR10A2_LE, RGB10A2_LE, Y410, A444, GBRA, AYUV, VUYA, RGBA, RBGA, ARGB, BGRA, ABGR, A422, A420, AV12, RGBP_F32LE, RGBP_F32BE, RGB_F32LE, RGB_F32BE, Y444_16LE, GBR_16LE, Y444_16BE, GBR_16BE, Y216_LE, Y216_BE, v216, P016_LE, P016_BE, Y444_12LE, GBR_12LE, Y444_12BE, GBR_12BE, I422_12LE, I422_12BE, Y212_LE, Y212_BE, I420_12LE, I420_12BE, P012_LE, P012_BE, RGBP_F16LE, RGBP_F16BE, RGB_F16LE, RGB_F16BE, Y444_10LE, GBR_10LE, Y444_10BE, GBR_10BE, BGR10x2_LE, RGB10x2_LE, r210, I422_10LE, I422_10BE, NV16_10LE40, NV16_10LE32, Y210, UYVP, v210, I420_10LE, I420_10BE, P010_10LE, NV12_10LE40, NV12_10LE32, P010_10BE, MT2110R, MT2110T, NV12_10BE_8L128, NV12_10LE40_4L4, Y444, BGRP, GBR, RGBP, NV24, v308, IYU2, RGBx, xRGB, BGRx, xBGR, RGB, BGR, Y42B, NV16, NV61, YUY2, YVYU, UYVY, VYUY, I420, YV12, NV12, NV21, NV12_16L32S, NV12_32L32, NV12_4L4, NV12_64Z32, NV12_8L128, Y41B, IYU1, YUV9, YVU9, BGR16, RGB16, BGR15, RGB15, RGB8P, GRAY_F32LE, GRAY_F32BE, GRAY16_LE, GRAY16_BE, GRAY_F16LE, GRAY_F16BE, GRAY10_LE16, GRAY10_LE32, GRAY8 }
width: [ 1, 2147483647 ]
height: [ 1, 2147483647 ]
framerate: [ 0/1, 2147483647/1 ]
Properties
beam-width
“beam-width” guint
Number of prefixes to keep during beam search
Flags : Read / Write
Default value : 10
decoding-method
“decoding-method” Ctc-text-tensor-dec-method *
CTC decoding method
Flags : Read / Write
Default value : beam-search (0)
dictionary-file
“dictionary-file” gchararray
Dictionary file with one UTF-8 token per line
Flags : Read / Write
Default value : NULL
implicit-space
“implicit-space” gboolean
Append an implicit space token after the dictionary entries
Flags : Read / Write
Default value : false
top-k
“top-k” guint
Number of token candidates to consider at each beam search step (0 = all)
Flags : Read / Write
Default value : 0
Named constants
Ctc-text-tensor-dec-method
Members
beam-search (0) – BeamSearch: Decode using CTC prefix beam search
greedy (1) – Greedy: Decode by selecting the most probable token at each timestep
The results of the search are