FervorCreative AI
Live Latest 04.10.26 · morning 107 tools tracked 359 workflows indexed 274 topics Hot: MiniMax H3, ACE-Step 1.5, fal

AI video models fake pixel art because nothing forces them onto a grid or a fixed set of colors, and two new MiniMax H3 add-ons plus a three-command cleanup pass (shrink, one palette for the whole clip, enlarge without smoothing) turn their output into sprite animation that holds up, at the cost of an experimental, license-limited model and a step you have to do by hand.

PixelTune MiniMax H3 LoRApixelart8x H3 LoRAMiniMax H3falffmpegvideo-genlora-finetuningcreative-workflowsprompt-craftopen-weights

Real Pixel Art From an AI Video Model: Draw Big, Snap to the Grid, Lock the Palette

Two new add-ons teach the MiniMax H3 video model to paint in clean blocks. Paired with a three-step cleanup, they produce sprite animation that survives a close look. Here is the full workflow, what it costs, and where it still cheats.

Ask almost any AI video model for "pixel art" and it hands you something that looks right from across the room. Lean in and the illusion falls apart. The "pixels" are different sizes. Edges are soft where they should be hard. And the colors shimmer, because every frame picked its own slightly different shade of the same blue.

That shimmer is the giveaway. Real sprite artists work on a fixed grid with a short, fixed list of colors, and they hold both for the whole animation. A video model has neither rule. It paints smooth images and imitates the look of blocks.

On October 3, two people separately posted add-ons that try to fix the first half of that problem for MiniMax H3, the open video model MiniMax posted in late July. They arrived about half an hour apart. Neither fixes the second half on its own. That part is a cleanup pass you can run in three commands, and it is where most of the quality comes from.

The two add-ons

An add-on here is a small file you attach to the main model to teach it one style. Both of these teach the same thing: paint in blocks.

pixelart8x aims for "one colour per cell on a fixed 8 px grid, a limited palette, and frame-by-frame sprite-style motion." It works from a fixed opening phrase. For animation, start your prompt with "pixel art animation, crisp 8x pixel grid, limited palette." then describe your scene, and the card's example ends with "Sound: none." It wants frame sizes in steps of 32 pixels.

PixelTune is stricter. You tell it the size of the "real" sprite canvas, say 128 by 128, and generate at four times that size. Every true pixel becomes a 4 by 4 block, which gives the cleanup step something exact to aim at. Its card spells out the canvas in the prompt itself: "Native pixel canvas: 128 by 128 pixels. Each native pixel is a 4 by 4 solid-color square; ..."

My position: PixelTune's idea is the right one, even though its file is earlier and rougher. Declaring the grid up front and then enforcing it afterward is how you get pixel art instead of a pixel-art filter. pixelart8x is the easier one to try today, because its 8-pixel blocks are big enough to survive a service that does not let you choose exact frame sizes.

Both carry the MiniMax H3 Community License, which includes territory restrictions. Read it before you sell anything you make with them.

Why the cleanup pass matters more than the model

Even a model trained on clean grids drifts. A block edge lands one pixel off. Two neighboring blocks blend. A shadow uses a color that appears in one frame and never again.

The cleanup does three things, and each one fixes a specific tell:

  1. Shrink to the true sprite size. If the model painted in 4 by 4 blocks, shrink the video to a quarter of its width and height. Each block collapses to one pixel. Soft edges and in-between shades get averaged into a single color per pixel.
  2. Pick one palette for the whole clip. Build one list of, say, 16 colors from every frame at once, then force every pixel in every frame onto that list with no dithering (the speckle pattern tools use to fake extra colors). This kills the shimmer.
  3. Enlarge without smoothing. Scale back up with nearest-neighbor, the mode that copies each pixel into a hard-edged block instead of blending. The result looks like it was drawn on a grid because now it was.

PixelTune's own card describes the same three moves for its official cleanup: snap the grid, use a shared palette for the whole clip "without dithering," and enlarge with nearest-neighbor. You can reproduce it with ffmpeg, the free command-line video tool.

Put this into practice

You can do this without a graphics card. fal runs MiniMax H3 with add-ons at its H3 text-to-video LoRA page, where you paste the add-on's file link into the LoRA (add-on) field. At 480p it costs $0.0625 per second, so a five-second test is about 31 cents.

Step 1: Generate. Use pixelart8x for your first try. Paste the file link (https://huggingface.co/JOYUNGI/pixelart8x-h3-lora/resolve/main/pixelart8x_h3_v3.safetensors) and set the strength to 1. Choose 480P and the 1:1 aspect ratio for a square sprite. Write the prompt as:

pixel art animation, crisp 8x pixel grid, limited palette. A small knight walks in place, cape swaying, on a flat dark background. Sound: none.

Keep the action small and the background plain. Sprite sheets live on simple, looping motion, and the add-on's training clips were mostly full-body characters on flat backgrounds.

One setting deserves attention. fal's prompt expansion, which rewrites your prompt to add detail, is on by default ("balanced"). I would switch it to "disabled" for this. Those opening trigger words are the whole point, and a rewrite can bury them.

Step 2: Check the frame size. Because fal picks the exact pixel size for you, find out what you got. ffprobe, which comes with ffmpeg, tells you:

ffprobe -v error -select_streams v:0 -show_entries stream=width,height -of csv=p=0 knight.mp4

Divide each number by your block size (8 for pixelart8x, 4 for PixelTune). If it does not divide evenly, crop a few pixels off before the next step so the grid lines up.

Step 3: Shrink. For an 8-pixel grid:

ffmpeg -i knight.mp4 -vf "scale=iw/8:ih/8:flags=area" -an small.mp4

The area mode averages each block into one color, which is what you want here. The -an drops the audio track. MP4 files need even widths and heights, so if your shrunken size comes out odd, crop two more pixels in step 2 or write straight to a GIF instead.

Step 4: Build one palette, then apply it.

ffmpeg -i small.mp4 -vf "palettegen=max_colors=16" palette.png
ffmpeg -i small.mp4 -i palette.png -lavfi "paletteuse=dither=none" small.gif

The first command looks at every frame and picks 16 colors for the whole clip. The second forces every frame onto them. Try 8, 16 and 32 colors and pick the one that reads best.

Step 5: Enlarge for display.

ffmpeg -i small.gif -vf "scale=iw*8:ih*8:flags=neighbor" knight_big.gif

Keep small.gif too. That tiny file is your real asset. A game engine or an editor like Aseprite can open it at true size, and you can touch up the frames that went wrong by hand.

If you have a big graphics card and want PixelTune instead, its card documents a tested setup in the DiffSynth-Studio toolkit: 50 steps, strength 1, and canvases in multiples of 8 generated at four times size. Then run the same cleanup with a block size of 4.

Where it breaks

The models are experiments. PixelTune is an early save from a training run planned to go five times longer, built from 153 clips, and its author writes that "static-region stability and exact pixel boundaries may be inconsistent." pixelart8x learned from 20 stills and 60 clips, mostly anime characters in military-style outfits on flat backgrounds. Ask for a forest scene and expect the style to wobble.

Neither card says it was tested on fal. PixelTune's settings come from a different toolkit. fal is the convenient route, not the proven one, so treat your first few renders as calibration.

The cleanup cannot invent intent. Shrinking averages whatever the model painted. If it drew a hand as a smudge, you get a one-pixel smudge. Bad frames still need a human with a pencil tool.

Motion is the hard part. Real sprite animation uses a handful of held poses at low frame rates. A video model gives you 24 frames per second of continuous motion. You may need to drop frames (keep every third, for example) to get the snappy, stepped feel of hand-made sprites, and the card does not test that.

The license follows the file. Everything rides on the MiniMax H3 Community License, with its territory limits. Check it before a client sees the result.

The habit worth keeping

The add-ons will improve or be replaced within weeks. The cleanup will not need to. Shrink to the true grid, lock one palette across the clip, enlarge without smoothing: those three moves turn any model's "pixel-art style" into actual pixel art, or show you exactly where it failed.

Run one five-second knight through it. Look at the small GIF at true size. If it reads as a sprite, you have a new way to block out a game's animations in an afternoon. If it does not, you will know whether the grid or the palette gave it away, and that tells you which add-on to try next.


Medium metadata

Title: Real Pixel Art From an AI Video Model: Draw Big, Snap to the Grid, Lock the Palette

Subtitle: Two new MiniMax H3 add-ons paint in clean blocks, and a three-command ffmpeg pass locks the grid and the colors. The full workflow, the cost, and where it still cheats.

Tags: Pixel Art, AI Video, Game Development, Animation, Generative AI

Estimated read time: 8 minutes