FervorCreative AI
Live Latest 30.09.26 · morning 86 tools tracked 304 workflows indexed 244 topics Hot: MiniMax H3, Qwen-Image-2.1, Ming-Image-0.1-Design

A new add-on for MiniMax H3 turns a written scene into a look-around 360 video with sound for about two dollars, and the craft lies less in the generation than in writing for a viewer who can face any direction and in the two small finishing steps that make a headset recognise the file.

MiniMax H3MiniMax H3 360 equirect LoRAfalSpatial Media Metadata Injectorvideo-gencreative-workflowslora-finetuningprompt-craft

Make a 360 VR Video With Sound From a Text Prompt, Then Watch It in a Headset

A normal video prompt describes what the camera sees. A 360 prompt has to describe what is behind you.

That sounds like a small difference until you try it. In a headset, the viewer can turn around. If your prompt only covered the waterfall in front of them, whatever the model invents behind their back is a guess, and people always look behind them. Plenty of 360 filmmakers learn this on their first shoot, usually by finding the crew in the back of the shot.

On September 29 (UTC), an add-on appeared on Hugging Face that lets the open video model MiniMax H3 make this kind of footage from a sentence: a full sphere around the viewer, with sound, ready to play on a Quest or on YouTube 360 after two short finishing steps. The author, who posts as rehan-fal, prices a 10-second 4K clip at $2.00 on fal's hosted service, taking roughly four to nine minutes to render.

The clever part is not the price. It is that the model card reads like a field guide for a new kind of writing, and it names the one setting that can ruin your results without any warning. Here is the whole workflow, from prompt to headset.

What the add-on does

A 360 video, laid flat, looks like a stretched world map. Straight ahead sits in the middle of the frame. Directly behind you is split between the far left and far right edges, which must match perfectly because in the headset they join. The sky and ground are stretched across the top and bottom. The technical name for this flat layout is "equirectangular," and you will see it in file names and settings.

Ordinary video models do not know this layout. Ask plain MiniMax H3 for a "360 panorama" and you get a nice wide shot that falls apart in a headset. The add-on teaches it the layout. It was trained on 181 real 360 clips, drawn from a public index of 360 videos and kept only after they passed checks for matching edges and properly stretched poles.

The author measured the result. The left and right edges of the frame, which need to match, line up at a score of 0.95 with the add-on, against 0.55 for plain H3 and 0.92 for the real training footage. In plain terms: the seam behind the viewer mostly disappears.

H3 generates sound with the picture, so you get rain, birds, traffic or a waterfall along with the image. That matters more in 360 than anywhere else, because sound is how you tell a viewer where to look.

Why the writing is the hard part

The card's prompt advice is short and specific. Start with the trigger word eqr360, then a fixed sentence that describes the layout (it was in every training caption, so leave it word for word), then your scene. And describe "what is ahead, to the sides, behind and overhead; the viewer can look everywhere."

Look at one of the author's own sample scenes, a rainy alley in Tokyo:

...glowing neon signs and red paper lanterns on both sides, steam rising from a tiny ramen stall straight ahead, people with clear umbrellas walking past, wet pavement reflecting pink and blue light, a train rumbling across a bridge behind, rain pattering, sizzling food and distant city sounds

Every direction gets something. Sides: lanterns. Ahead: the ramen stall. Behind: the train. Below: reflecting pavement. And the sounds match the places: the sizzle belongs to the stall ahead, the rumble to the train behind.

That is how a set designer thinks about a room, not how a cinematographer thinks about a frame. If you write the way you normally prompt a video, with one subject and a camera move, you will get a strong front and a vague, smeared back half.

Put this into practice

You need a fal account with credit, a free tool called ffmpeg for the conversion, and Google's free Spatial Media Metadata Injector. A headset helps but is not required: YouTube and most desktop 360 players let you drag around the view with a mouse.

1. Write the scene in four directions. Before you touch the tool, jot one line each for ahead, left and right, behind, and overhead or underfoot. Add two or three sounds and say where they come from. Keep the scene simple and wide. Open landscapes, rooms and streets work better than a close-up of a person.

2. Assemble the prompt. Trigger word, then the layout sentence copied exactly from the model card, then your scene. Keep the whole thing in plain descriptive sentences.

3. Run a cheap preview first. On fal's MiniMax H3 text-to-video endpoint with add-ons, attach the file at strength 0.75, set the shape to 21:9, and turn prompt expansion off so the service does not rewrite your careful directions. Start at 768P, which the card says renders in about a minute. Note the seed.

4. Pass the direct file link, not the repo name. This is the setting that bites. The card says that when the add-on was loaded by its Hugging Face repo name on fal, the same seed produced a different video, while the direct link to the file reproduced the trained result exactly. Use the full link ending in h3-360-equirect-lora-v1.safetensors.

5. Render the keeper at 2K or 4K with the same seed. The higher tiers upscale the 768P pass, so your preview is a faithful guide to the final scene. A sphere spreads pixels thin, so go as high as your budget allows. At 4K the card puts a 10-second clip at $2.00.

6. Keep it short. The card says 5 to 15 seconds all hold the layout, with 5 seconds giving the cleanest seam. Longer renders can darken in their last few seconds, so for a 10-second clip, plan to keep the first seven and a half.

7. Convert the shape. The output is 21:9. A 360 player expects exactly 2:1. The card gives a one-line ffmpeg command that resizes a 4K result to 4096 by 2048 and keeps the audio untouched. Copy it from the card rather than retyping it.

8. Tag it as 360. Players decide whether a file is 360 by reading a hidden label inside it. Google's Spatial Media Metadata Injector adds that label. Mark it as a spherical, mono video. YouTube's own help page points to this tool (or Adobe Premiere) for 360 uploads and warns that 360 playback can take up to an hour to switch on after upload.

9. Watch it where it will be seen. Copy the file to a Quest, a player such as DeoVR or Skybox, or upload it to YouTube. Turn around. Look up. Check the seam behind you.

If the seam shows, the author's source repository includes a packaging step that softens it and can make short clips loop smoothly, which helps because headset players often loop short files. Every sample on the card was run through it.

Where it breaks

The seam is reduced, not gone. The card says a faint line can remain directly behind the viewer on busy scenes. H3 was not designed to wrap around, so the two edges can disagree slightly. Calm, open scenes hide it best.

It is flat, not 3D. This is a mono sphere. Nothing will feel close enough to touch. The same author has a separate VR180 add-on that produces stereo depth, but only for the half of the world in front of you.

The training footage leaves fingerprints. The 181 clips skew toward handheld and first-person uploads, so the card warns you may see people close to the camera and a slightly high horizon. If a stranger appears at your elbow, rerun with a new seed and push the scene's people further away in the prompt.

Detail runs thin. A 4K 360 frame gives you about 11 pixels per degree of view, per the card. That is fine for landscapes and atmosphere. It is soft for anything you want the viewer to read, such as signs or text.

The licence comes with it. The add-on inherits the MiniMax Community License from H3, including its restrictions. Read it before you put a clip in paid work. The training clips came from YouTube and are not redistributed, which the author explains: the add-on learns a layout, not the content of those videos.

Headset players are picky. If your file opens as a flat stretched rectangle, the metadata step did not take. Run the injector again before you blame the model.

Why this one is worth an evening

For a long time, making a 360 scene meant a special camera rig, stitching software and a location you could stand in. Generated 360 used to be a curiosity: pretty from the front, broken behind you. This add-on stands out because it publishes its seam measurements, names its known failures and shows you how to package the result.

That makes it useful beyond VR fans. A set designer can mock up a room a client can stand inside. A musician can make a looping world behind a track. A game team can sketch a location's mood, sound included, before anyone builds it. None of these need perfection. They need a sense of place.

Pick one place you know well, write it in four directions, and spend two dollars finding out what the model makes of it. Then put on the headset, turn around, and see whether the back of your world holds up. If it does, you have a new sketchbook. If it does not, you will know exactly which line of the prompt to rewrite.


Medium metadata

Title: Make a 360 VR Video With Sound From a Text Prompt, Then Watch It in a Headset

Subtitle: A new add-on for MiniMax H3 turns a written scene into a look-around 360 clip for about $2. The trick is writing for a viewer who can turn around.

Tags: Virtual Reality, AI Video, 360 Video, Filmmaking, Generative AI

Estimated read time: 8 minutes