The SIA Challange V1 is based on text data recommendation. The searching or queying for an item is dependent on description, title, additional information etc of an item. Entire text is represented in the form of vector, that uses "Bidirectional Transformer Model" as in refrence note 2.
The architecture of V1 is represented in the figure below.

(Details on Pitchbook)
Version_2 is interactive and is based on computer vision. User can upload the image or use live streaming from the mobile camera or webcam, the captured image or streaming video (<=30 fps), will under object detection with the given input. (Demo available under the image/demo) directory.

