← 返回信息流

Meta AI推文

Muse Spark 1.2 支持广泛多模态任务

模型AI评分:70/100

Muse Spark 1.2 支持从视觉到代码、感知到行动等多种多模态任务,并具备强大的音视频理解能力,适用于企业视频密集型工作流。官方发布了新的评估和演示,展示模型在视觉理解和推理方面的能力,包括引导机器人在非结构化环境中找到橡皮鸭的示例。

译文

Muse Spark 1.2 支持广泛的多模态任务,从将视觉内容转化为可运行的代码,到将感知转化为物理动作。它还带来了强大的音视频理解能力,以支持企业实际应用中常见的视频密集型工作流。今天,我们分享新的评测和演示,展示该模型在视觉理解与推理能力上的广度。让我们从一个演示开始,看看 Muse Spark 如何解析多模态观察结果并调用工具,引导机器人在非结构化环境中导航,找到一只橡皮鸭。🧵👇

M
Meta AI

@AIatMeta

Muse Spark 1.2 supports a broad range of multimodal tasks, from turning visuals into working code to translating perception into physical action. It also brings robust audio-visual understanding to enable video-heavy workflows common in real-world enterprise use. Today, we’re sharing new evals and demos that illustrate the breadth of the model’s visual understanding and reasoning capabilities. Let’s start with a demo that shows how Muse Spark parses multimodal observations and calls tools to guide a robot to navigate in an unstructured environment to find a rubber duck. 🧵👇

阅读原文