How Lenta Tech Searched for Price Tags on Video from a Moving Robot: Results of the Video Analytics Hackathon
Ultralytics
PaddleOCR (Baidu)
OpenCV
EasyOCR (Jaided AI)
Tesseract OCR (Google)
The Lenta Tech team (IT brand of the Lenta Group) shared their experience of conducting a hackathon on recognizing price tags from video captured by a moving robot. The article describes the solution pipeline: from detecting price tags using YOLO to extracting data via OCR, QR codes, and a catalog. Key takeaways: quality is achieved not by a single model but by proper handling of frame sequences (tracking, selecting the best frame) and combining OCR, QR, and catalog.
The Lenta Tech team, the IT brand of the Lenta Group, held a hackathon in 2025 called Lenta Tech Life Hack, focused on recognizing price tags from video footage captured by a robot moving along shelves in supermarkets. The task proved more challenging than simple rectangle detection and OCR, due to motion blur, glare, varying angles, partial occlusions, and a wide variety of price tag formats. Participants were provided with video recordings from different store zones (alcohol, dairy, honey, etc.) with both stops and continuous motion; the output had to be a CSV file with fields: product name, prices, barcode, SKU, print date, display zone, coordinates, timestamp, and data from QR codes. Restrictions included using only local open-source models, no external online services or cloud APIs, and optimization for rknn int8 was encouraged. The evaluation metric: the proportion of price tags recognized with at least 80% accuracy. The first round showed that the problem lay not only in detection but also in correct reading and data structuring. Successful solutions used a cascaded pipeline: tracking and best-frame selection (sharpest, largest, most front-facing), detection (mostly YOLO/Ultralytics), rectification and super-resolution, zonal OCR (PaddleOCR, EasyOCR, Tesseract) with field validation via regular expressions and business rules, and data recovery from QR/barcodes and a catalog with fuzzy matching. In the final stage, average quality improved by 4.47 percentage points, with maximum gains of +21.27 and +19.48 percentage points, achieved through tracking, best-frame scoring, QR recovery, catalog search, sanity checks, and Docker packaging. The winning approach was a cascaded architecture rather than a single model. Stacks that performed well included: YOLO/Ultralytics, OpenCV, PaddleOCR, EasyOCR, Tesseract, QR/barcodes, catalog/SKU/fuzzy search, tracking, and best-frame selection. The approach 'find a price tag and then figure it out somehow' did not work — without accounting for the price tag structure and repeated frames.
Source: Habr — хаб ИИ —
original
