Manga Tone Rendering LoRA for MiniMax H3 v2.0

MiniMax H3 LORA
Download

A LoRA that teaches MiniMax H3 (video + audio joint DiT) the manga

screentone rendering style with motion: monochrome ink rendering, proper

tone shading, and tone that holds up under dynamic camera movement and action.

Usage

- Trigger phrase: manga-tone rendering

- Recommended strength: 1.0 (usable 0.8 - 1.2)

- Standalone: res_multistep / simple, 14 steps

- With turbo LoRA: euler / beta, 8 steps (ref2v turbo 8-step) or 4 steps

(fl2v turbo 4-step), audio sigma shift 5

- Checkpoints: int8 and mixed int4/int8 pruned convrot conversions

- Load through any H3 LoRA loader (H3LoraStack / LoraLoaderModelOnly)

Example prompt:

manga-tone rendering, a samurai draws his sword in a bamboo grove,

dramatic low angle, speed lines

What it does

Without the LoRA, H3 renders "manga" prompts as grayscale anime with flat

shading and little motion. With v2, outputs are monochrome with ink/midtone

statistics inside the trained manga distribution, and motion is preserved

under action prompts (measured +35% inter-frame motion vs v1 on a

motion-heavy prompt, tone metrics unchanged).

Training

All training data is synthetic; no published or third-party artwork was used:

- 714 static 5-frame crop clips from 58 AI-generated fictional manga pages,

two crop scales

- 80 animated clips: still crops animated by H3 itself (v1 @0.7 + turbo),

then cleaned frame-by-frame with a deterministic monochrome render pass

- 6 windowed clips from two action reference videos produced by feeding the

same fictional manga pages to Google Omni (x4 repeat)

- LoRA rank 32 / alpha 32 on 208 projections, flow matching on the video

stream, lr 5e-5 constant with warmup, 2000 steps, bf16, gradient

checkpointing; v2 initialized from v1

- Audio stream masked in the loss; audio generation unchanged to first

approximation but not guaranteed

Notes

- Tone dot frequency is learned relative to page scale (renders ~9 px period

at 960 px canvas width).

- Abstract motion-graphics tropes (layer decomposition, graphic animation)

remain out of distribution for the base model; action/cinematic prompts

play to this LoRA's strengths.