Amazon's New Toolkit Turns a Sentence Into a Short Video
Type a description, get an image, then watch it move — no camera needed.
Turn a text prompt into an image, then animate that image into a short video — on Amazon SageMaker AI. You deploy two endpoints from the same AWS vLLM-Omni Deep Learning Container (DLC): a real-time endpoint for FLUX.2-klein-4B image generation and an asynchronous endpoint for Wan2.1-VACE-1.3B video generation. You send the prompt to the image endpoint and receive a base64-encoded PNG; the workflow resizes it, converts it to a compact JPEG data URL, and passes it along with a motion prompt to the video endpoint, which stores the MP4 in Amazon S3 for retrieval. The sample uses fixed ml.g6.xlarge and ml.g6e.xlarge instance types and includes a command-line workflow and an optional Streamlit application.
- Two AI models work as a team: one draws a picture from your words, the second makes that picture move into a short video.
- Amazon's guide is free to read, but running it costs money — you rent graphics chips by the hour, roughly a dollar or two each.
- This is a developer recipe, not an app you can download, so ordinary users will see it arrive later inside other companies' products.
Why It Matters
Cheap, do-it-yourself video generation means small teams and solo creators can soon make video ads and clips without a camera crew.