Chinese Model SenseNova U1.5-Lite-Preview Goes Open Source with Native 4K Output
SenseTime
SenseTime has previewed and open-sourced SenseNova U1.5-Lite-Preview, a lightweight native unified multimodal model that can generate 4K images directly. It supports complex prompts, reference images, multi-image composition, and precise editing, marking a step forward in AI image generation and editing capabilities.
SenseTime released SenseNova U1.5-Lite-Preview, an early preview of a lightweight native unified multimodal model, to the open-source community. Built on the self-developed NEO-Unify architecture, it models language, visual semantics, and pixel generation within a single framework, covering visual understanding, reasoning, generation, and editing. Despite its small 8B-MoT parameter size, it can handle long, multi-constraint, hierarchical, and structured visual instructions, making it adept at creating high-information-density content like posters and infographics. The model also supports post-generation editing, including reference images, multi-image composition, and precise local edits using red boxes, coordinates, and markers. It demonstrated the ability to replicate and adapt design frameworks, such as converting a movie-themed infographic into a postal mail journey, and to modify single characters like changing '迎' into '命' by reorganizing visual elements. The model can also combine multiple reference images, transforming a real photo into a hand-drawn recipe, or a portrait into a one-page resume. SenseNova re-engineered the generation head to push native resolution to 4K, reducing grid artifacts and improving texture detail. Performance improved on benchmarks: Qwen-Image-Bench score rose from 47.14 to 55.20, ImgEdit-Bench from 3.90 to 4.37, and GEdit-Bench English from 7.47 to 8.17, Chinese from 7.42 to 8.05, with WeEdit Overall Average at 6.44. The official version of U1.5 is expected to be open-sourced soon. Links are provided on GitHub, Hugging Face, and ModelScope.
- Abbreviations
- MoT = Mixture of Tokens — Смесь токенов
- B = Billion — Миллиард
Source: QbitAI 量子位 —
original
