Research author
Paola Cascante-Bonilla
2 AI research papers in the One9Founders library, with summaries and links to original sources.
Papers by Paola Cascante-Bonilla
SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis
Kathakoli Sengupta, Kai Ao, Paola Cascante-Bonilla
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders
Shang-Jui Ray Kuo, Paola Cascante-Bonilla
This paper investigates whether state space models (SSMs) can replace the standard Vision Transformers used in vision-language models (VLMs) that combine images with text. Through systematic testing, the researchers found that SSM-based vision encoders perform comparably to or better than Vision Transformers, especially when trained on detection or segmentation tasks, while using fewer parameters. The study challenges the assumption that larger models or higher image classification accuracy automatically lead to better VLM performance and proposes stabilization techniques to improve robustness.