Research author
Dong Zhuo
2 AI research papers in the One9Founders library, with summaries and links to original sources.
Papers by Dong Zhuo
SM4RT: Learning Structured Motion Geometry for 4D Reconstruction
Shing Ho J. Lin, Wenzhao Zheng, Dong Zhuo, et al.
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
Dong Zhuo, Wenzhao Zheng, Sicheng Zuo, et al.
DriveTok is a new tool that converts multi-camera driving scenes into efficient digital 'tokens' (compressed representations) that capture semantic meaning, depth, and 3D spatial information all at once. Unlike existing tokenizers designed for single images, DriveTok is specifically built for autonomous vehicles with multiple cameras, using advanced attention mechanisms to ensure consistency across different camera views. The tokens can be used for various driving-related tasks like reconstructing images, segmenting objects, predicting depth, and understanding 3D space around the vehicle.