← Tags
#VLM (9)
Microsoft Mage-VL: 4B Streaming VLM for Real-Time Live Video Understanding
视觉语言模型VLM流式理解Twitter
Open-set object detection recognize unseen objects
感知Perception开放集Open-set检测
OMY + Gemini + GraspNet: VLM-Driven Smart Pick-and-Place
抓取放置智能抓取VLM场景理解6-DoF位姿
Astribot releases Lumo-2 World-Action Model
世界动作模型WAMLumo-2AstribotQwen
TOPReward zero-training robot reward model
VLAReward ModelVLM零训练ZeroTraining
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
视觉语言导航自感知推理VLMVLN具身智能
3D-Aware VLMs with Implicit and Explicit Geometries
3D感知3D-Aware视觉语言模型VLM几何
Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting
第一人称egocentric手部姿态hand poseVLM
Ludi₀.₁: An Agentic System for Socially Intelligent Robots
社交智能agentic systemLudo RoboticsVLM人机交互