开年王炸！OpenAI发布文本转视频模型Sora，有亿点震撼！

本文主要是介绍开年王炸！OpenAI发布文本转视频模型Sora，有亿点震撼！，希望对大家解决编程问题提供一定的参考价值，需要的开发者们随着小编来一起学习吧！

大家好，我是木易，一个持续关注AI领域的互联网技术产品经理，国内Top2本科，美国Top10 CS研究生，MBA。我坚信AI是普通人变强的“外挂”，所以创建了“AI信息Gap”这个公众号，专注于分享AI全维度知识，包括但不限于AI科普，AI工具测评，AI效率提升，AI行业洞察。关注我，AI之路不迷路，2024谷歌一起变强。

一些结论

Sora是OpenAI开发的文本转视频AI模型，可根据文本创建真实和富有想象力的视频场景。

Sora旨在理解和模拟物理世界的运动，解决现实世界互动问题。

该模型能生成长达一分钟的高质量视频，忠实反映用户指令。

Sora能构造包含多角色和动作的复杂场景，深刻理解物理世界。

通过扩散模型和变压器架构，Sora精确解读文本提示，生成生动情感的角色。

Sora利用补丁表示和DALL·E 3的重述技术，提高文本到视频的忠诚度。

Sora的开发标志着向实现AGI的重要步骤，模拟真实世界互动。

OpenAI采取多项安全措施，包括对抗测试和误导内容检测，确保Sora的安全使用。

Sora生成视频展示（来自OpenAI官方）

所有展示的Sora视频均未经修改，直接展现其生成能力。

东京霓虹灯下，一位自信女性的夜晚漫步

原提示词：A stylish woman walks down a Tokyo street filled with warm glowing neon and animated city signage. She wears a black leather jacket, a long red dress, and black boots, and carries a black purse. She wears sunglasses and red lipstick. She walks confidently and casually. The street is damp and reflective, creating a mirror effect of the colorful lights. Many pedestrians walk about.

好奇小怪物与融化蜡烛的温馨邂逅

原提示词：Animated scene features a close-up of a short fluffy monster kneeling beside a melting red candle. The art style is 3D and realistic, with a focus on lighting and texture. The mood of the painting is one of wonder and curiosity, as the monster gazes at the flame with wide eyes and open mouth. Its pose and expression convey a sense of innocence and playfulness, as if it is exploring the world around it for the first time. The use of warm colors and dramatic lighting further enhances the cozy atmosphere of the image.

纸艺珊瑚礁中的彩色海洋世界

原提示词：A gorgeously rendered papercraft world of a coral reef, rife with colorful fish and sea creatures.

穿越盐沙漠的30岁太空人冒险电影预告

原提示词：A movie trailer featuring the adventures of the 30 year old space man wearing a red wool knitted motorcycle helmet, blue sky, salt desert, cinematic style, shot on 35mm film, vivid colors.

雪地中巨大猛犸象的壮丽征途

原提示词：Several giant wooly mammoths approach treading through a snowy meadow, their long wooly fur lightly blows in the wind as they walk, snow covered trees and dramatic snow capped mountains in the distance, mid afternoon light with wispy clouds and a sun high in the distance creates a warm glow, the low camera view is stunning capturing the large furry mammal with beautiful photography, depth of field.

雪中东京，樱花与雪花共舞的城市风光

原提示词：“Beautiful, snowy Tokyo city is bustling. The camera moves through the bustling city street, following several people enjoying the beautiful snowy weather and shopping at nearby stalls. Gorgeous sakura petals are flying through the wind along with snowflakes.”

OpenAI正式发布Sora

Sora是OpenAI开发的一款AI模型，它能够根据文本指令创建真实和充满想象力的视频。其设计目标是让AI学会理解并模拟物理世界中的运动，从而帮助人们解决需要与现实世界互动的问题。Sora的出色之处在于它能生成长达一分钟的视频，同时确保视频的视觉质量以及对用户指令的忠实遵循。

Sora具备生成包含多角色、特定动作类型和精确主题及背景细节的复杂场景的能力。这表明该模型不仅理解用户提示中的请求内容，还理解这些内容在物理世界中是如何存在的。Sora能够精确解读文本提示，并生成表情生动、情感丰富的角色，同时在单个视频中创造多个镜头，准确保持角色和视觉风格的连贯性。

技术上，Sora是基于扩散模型，从类似静态噪声的视频开始，通过多个步骤逐步转换，去除噪声生成视频。它采用了与GPT类似的变压器架构，提高了扩展性能，并将视频和图像表示为称为“补丁”的小型数据单元集合，这类似于GPT中的令牌。借鉴了DALL·E和GPT的研究，Sora使用了DALL·E 3的重述技术，能更忠实地遵循用户的文本指令。除了能从文本指令生成视频外，Sora还能从现有静态图像生成视频，动画化图像内容，细致入微。

为了确保安全性，OpenAI在将Sora集成到其产品前，计划采取多项重要安全措施。这包括与领域专家合作进行对抗测试，他们是在误导信息、仇恨内容和偏见等方面的专家。OpenAI还在开发工具帮助检测误导性内容，包括一种能识别视频是否由Sora生成的分类器。计划未来引入C2PA元数据，并利用为DALL·E 3构建的现有安全方法。同时，OpenAI将与全球政策制定者、教育者和艺术家合作，了解他们的关切，并识别这项技术的积极用例。