RelateAnything: Real-Time Open-Vocabulary Relation Prediction

Loading video
Loading videoRelateAnything predicts ranked relations between any regions from any detector or segmenter at about 20 ms per frame with only 53M parameters. The predicate vocabulary is supplied at inference, so unseen relations work without retraining; it was trained on 474K images and 4.3M+ relations covering 10,102 predicates, and runs as one ONNX graph even in the browser. Code: github.com/Maelic/RelateAnything.
Scene Graph GenerationOpen-vocabularyReal-TimeRelation PredictionVisual perceptionRelateAnythingScene graphs
Category: perception
Author: @tokufxug
Date: 2026-09-16T00:00:00
Duration: 30.0s
Reference: https://arxiv.org/abs/2609.12552





