Cosmos 3: Omnimodal World Models for Physical AI

摘要:

我们引入Cosmos 3,一个全模态世界模型家族,模型设计在一个统一的MoT(Mixture-of-Transformers)架构中,联合处理和生成语言、图像、视频、音频以及行动序列。通过支持高柔性的输入输出配置,Cosmos 3无缝的统一了物理 AI的关键模态,在一个单一框架中包含了视觉语言模型,视频生成模型,世界仿真器,和世界行动模型。

Our evaluation demonstrates that Cosmos 3 establishes
a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstrating
omnimodal world models as scalable, general-purpose backbones for embodied agents. Our post-trained
Cosmos 3 models were ranked as the best open-source Text-to-Image and Image-to-Video models by Artificial
Analysis, and the best policy model by RoboArena at the time the technical report was written. To
accelerate open research and deployment in Physical AI, we make our code, model checkpoints, curated
synthetic datasets, and evaluation benchmark available under the Linux Foundation’s OpenMDW-1.1
License at github.com/nvidia/cosmos and huggingface.co/collections/nvidia/cosmos3 . The project
website is available at research.nvidia.com/labs/cosmos-lab/cosmos3 .

评论
添加红包

请填写红包祝福语或标题

红包个数最小为10个

红包金额最低5元

当前余额3.43前往充值 >
需支付:10.00
成就一亿技术人!
领取后你会自动成为博主和红包主的粉丝 规则
hope_wisdom
发出的红包
实付
使用余额支付
点击重新获取
扫码支付
钱包余额 0

抵扣说明:

1.余额是钱包充值的虚拟货币,按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载,可以购买VIP、付费专栏及课程。

余额充值