SenseTime open-sources 8B multimodal model with native 4K image output
I'm LongbridgeAI, I can summarize articles.SenseTime has open-sourced SenseNova U1.5 Lite, an 8-billion-parameter multimodal model combining visual understanding, image generation, and editing. It supports native 4K output and handles complex constraints like subjects, spatial relationships, and layouts. The lightweight version improves identity preservation and offers controls such as bounding boxes and reference images. Available on GitHub, Hugging Face, and ModelScope.
SenseTime has open-sourced SenseNova U1.5 Lite, an 8-billion-parameter multimodal model designed to combine visual understanding, image generation and editing in one system. The model supports native 4K image output and is designed to handle constraints involving subjects, counts, spatial relationships, text, layouts and visual styles.
The company says the lightweight model improves identity preservation and spatial structure during editing, while adding controls such as bounding boxes, visual markers and multiple reference images. SenseNova U1.5 Lite is available through GitHub, Hugging Face and ModelScope. [[IT Home, in Chinese]
