Research author
Jiwen Lu
7 AI research papers in the One9Founders library, with summaries and links to original sources.
Papers by Jiwen Lu
SM4RT: Learning Structured Motion Geometry for 4D Reconstruction
Shing Ho J. Lin, Wenzhao Zheng, Dong Zhuo, et al.
Point Cloud Diffusion with Global and Local Reconstruction for Instance-Level 3D Anomaly Detection
Linchun Wu, Qin Zou, Jiwen Lu, et al.
BAMI: Training-Free Bias Mitigation in GUI Grounding
Borui Zhang, Bo Zhang, Bo Wang, et al.
Attention at Rest Stays at Rest: Breaking Visual Inertia for Cognitive Hallucination Mitigation
Boyang Gong, Yu Zheng, Fanye Kong, et al.
DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale
Sicheng Zuo, Zixun Xie, Wenzhao Zheng, et al.
Vega: Learning to Drive with Natural Language Instructions
Sicheng Zuo, Yuxuan Li, Wenzhao Zheng, et al.
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
Dong Zhuo, Wenzhao Zheng, Sicheng Zuo, et al.
DriveTok is a new tool that converts multi-camera driving scenes into efficient digital 'tokens' (compressed representations) that capture semantic meaning, depth, and 3D spatial information all at once. Unlike existing tokenizers designed for single images, DriveTok is specifically built for autonomous vehicles with multiple cameras, using advanced attention mechanisms to ensure consistency across different camera views. The tokens can be used for various driving-related tasks like reconstructing images, segmenting objects, predicting depth, and understanding 3D space around the vehicle.