Paper
RayRoPE: Projective Ray Position Encoding for Multi-View Attention
Apple and CMU jointly propose RayRoPE, a position encoding scheme for multi-view Transformers. It is based on associated rays and utilizes 3D points predicted along the rays for geometry-aware encoding. By computing projected coordinates in the query frame, it achieves SE(3) invariance and can analytically compute expected position encodings when predicted points are inaccurate. In novel view synthesis and stereo depth estimation tasks, RayRoPE achieves a relative 15% improvement in LPIPS on the CO3D dataset and can seamlessly fuse RGB-D inputs.
Read the original (opens in a new tab)
News stream data aggregated by AI HOT