EN Submit a tool
Paper

RayRoPE: Projective Ray Position Encoding for Multi-View Attention

Published: Source: Apple Machine Learning Research (RSS)

ShareXFacebookTelegramWhatsApp

Apple and CMU jointly propose RayRoPE, a position encoding scheme for multi-view Transformers. It is based on associated rays and utilizes 3D points predicted along the rays for geometry-aware encoding. By computing projected coordinates in the query frame, it achieves SE(3) invariance and can analytically compute expected position encodings when predicted points are inaccurate. In novel view synthesis and stereo depth estimation tasks, RayRoPE achieves a relative 15% improvement in LPIPS on the CO3D dataset and can seamlessly fuse RGB-D inputs.

Read the original (opens in a new tab)

News stream data aggregated by AI HOT

Related newsLatest in this category
· X: Elvis Saravia (@omarsar0, DAIR.AI)
· Hacker News Hot (buzzing.cc Chinese translation)
· The Decoder: AI News (RSS)
· X: Rohan Paul (@rohanpaul_ai)
· X: Rohan Paul (@rohanpaul_ai)