lovart能将单张高清图片转化为具备镜头语言与叙事逻辑的20秒电影感视频。需上传≥1080p无遮挡原图,固定选用nano banana pro模型,再用指定提示词生成25帧分镜脚本,最后通过可灵o1按帧生成并导出mp4成片。
☞☞☞AI 智能聊天, 问答助手, AI 智能搜索, 多模态理解力帮你轻松跨越从0到1的创作门槛☜☜☜

你想用一张现有图片生成一段连贯、有电影感的视频,不是简单动效而是具备镜头语言和叙事逻辑的成片,Lovart能直接在画布上完成从静态到动态的全流程转化。
准备原始图片与模型选择
打开 lovart.ai,登录后点击「新建工程」→ 上传你选定的参考图(JPG/PNG,建议分辨率不低于1080p);【必须确保图片中所有主体清晰可见、无严重遮挡或模糊】。上传完成后,在左侧模型栏选择「Nano Banana Pro」——这是目前Lovart中图像一致性最强、细节还原度最高的基础模型,专为后续视频生成提供稳定锚点。
这一步不能跳过或换用其他模型,否则可灵O1在生成视频时会因主体特征漂移导致人物/物体变形、光影断裂。
生成结构化分镜脚本
在对话框中粘贴以下完整提示词(无需修改句式,直接复制):
You are an award-winning trailer director+cinematographer+storyboard artist. Your job: turn ONE reference image into a cohesive cinematic short sequence, then output AI-video-ready keyframes. User provides: one reference image (image). 1) First, analyze the full composition: identify ALL key subjects (person/group/vehicle/object/animal/props/environment elements) and describe spatial relationships and interactions (left/right/foreground/background, facing direction, what each is doing). 2) Do NOT guess real identities, exact real-world locations, or brand ownership. Stick to visible facts. Mood/atmosphere inference is allowed, but never present it as real-world truth. 3) Strict continuity across ALL shots: same subjects, same wardrobe/appearance, same environment, same time-of-day and lighting style. Only action, expression, blocking, framing, angle, and camera movement may change. 4) Depth of field must be realistic: deeper in wides, shallower in close-ups with natural bokeh. Keep ONE consistent cinematic color grade across the entire sequence. 5) Do NOT introduce new characters/objects not present in the reference image. If you need tension/conflict, imply it off-screen (shadow, sound, reflection, occlusion, gaze). Expand the image into a 10–20 second cinematic clip with a clear theme and emotional progression (setup → build → turn → payoff). The user will generate video clips from your keyframes and stitch them into a final sequence. Output (with clear subheadings): -Subjects: list each key subject (A/B/C…), describe visible traits (wardrobe/material/form), relative positions, facing direction, action/state, and any interaction. -Environment & Lighting: interior/exterior, spatial layout, background elements, ground/walls/materials, lighting direction and quality...
发送后等待约30秒,Lovart会自动输出包含「场景分解」「主题与故事」「电影手法」「25个关键帧列表」的完整分镜文档。这个文档是后续视频生成的唯一依据,不可手动删减或合并条目。
调用可灵O1生成视频片段
方法一:全图驱动生成
选中画布上的原始图片 → 拖入右侧对话框 → 输入指令:“用可灵O1模型,按刚才生成的25帧分镜脚本生成视频,每帧时长0.8秒,总时长约20秒,保持原始构图比例与色彩基调”。点击生成,约2~4分钟完成首版视频。
方法二:局部强化生成(适用于某段镜头质感不足)
在已生成的视频预览中暂停到第7帧画面 → 点击「Touch Edit」按钮 → 用鼠标圈出人物面部区域 → 输入:“增强皮肤纹理真实感,保留原光照方向与阴影结构,不改变表情与朝向” → 再次生成该片段并自动替换原帧。
【注意:两次生成必须使用同一账号下的积分余额,跨账号无法继承分镜上下文】
导出与格式设置
视频生成完毕后,点击右上角「导出」按钮 → 弹窗中选择「MP4|H.264|1080p|30fps|无水印」→ 勾选「保留原始音频轨道(如有)」→ 点击「开始导出」。
导出进度条走完即完成,文件将自动下载至本地默认下载目录,命名规则为“lovart_日期_时间.mp4”,无需重命名即可直接用于发布或剪辑嵌入。











