* Equal Contribution Corresponding Author

2026

RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos
RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos

Zhenhua Wu*, Yun Pang*, Mingkun Chang*, Yuwei Ning, Liangzhi Wang, Yi Xiao, Guanbin Li

arXiv preprint

RealityBridge is a structure-preserving and asset-aware Sim-to-Real framework for improving edited 3DGS driving videos, targeting artifact removal, realism, and temporal consistency.

RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos

Zhenhua Wu*, Yun Pang*, Mingkun Chang*, Yuwei Ning, Liangzhi Wang, Yi Xiao, Guanbin Li

arXiv preprint

RealityBridge is a structure-preserving and asset-aware Sim-to-Real framework for improving edited 3DGS driving videos, targeting artifact removal, realism, and temporal consistency.

CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation

Chengzhuo Tong*, Mingkun Chang*, Shenglong Zhang, Yuran Wang, Cheng Liang, Zhizheng Zhao, Ruichuan An, Bohan Zeng, Yang Shi, Yifan Dai, Ziming Zhao, Guanbin Li, Pengfei Wan, Yuanxing Zhang, Wentao Zhang

ICML 2026

CoF-T2I studies how video models can perform frame-by-frame visual reasoning for text-to-image generation, using intermediate frames as explicit reasoning steps toward the final image.

CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation

Chengzhuo Tong*, Mingkun Chang*, Shenglong Zhang, Yuran Wang, Cheng Liang, Zhizheng Zhao, Ruichuan An, Bohan Zeng, Yang Shi, Yifan Dai, Ziming Zhao, Guanbin Li, Pengfei Wan, Yuanxing Zhang, Wentao Zhang

ICML 2026

CoF-T2I studies how video models can perform frame-by-frame visual reasoning for text-to-image generation, using intermediate frames as explicit reasoning steps toward the final image.