Nodes/ComfyUI-easygoing-nodes/Model Scale ERNIE Image
ComfyUI Node

Model Scale ERNIE Image

Zero out or amplify specific ERNIE-Image layers, down to the attention heads

By easygoing0114·Created 12 months ago·Updated 5 days ago· 5
Model Scale ERNIE Image
  • model
  • MODEL
x_embedder.1.00
text_proj.1.00
time_embedding.1.00
layers.0.1.00
layers.0.self_attention.1.00
layers.0.self_attention.to_q.1.00
layers.0.self_attention.to_k.1.00
layers.0.self_attention.to_v.1.00
layers.0.self_attention.to_out.1.00
layers.0.self_attention.norm_q.1.00
layers.0.self_attention.norm_k.1.00
layers.0.mlp.1.00
layers.0.mlp.gate_proj.1.00
layers.0.mlp.up_proj.1.00
layers.0.mlp.linear_fc2.1.00
layers.0.adaLN_mlp_ln.1.00
layers.0.adaLN_sa_ln.1.00
layers.1.1.00
layers.1.self_attention.1.00
layers.1.self_attention.to_q.1.00
layers.1.self_attention.to_k.1.00
layers.1.self_attention.to_v.1.00
layers.1.self_attention.to_out.1.00
layers.1.self_attention.norm_q.1.00
layers.1.self_attention.norm_k.1.00
layers.1.mlp.1.00
layers.1.mlp.gate_proj.1.00
layers.1.mlp.up_proj.1.00
layers.1.mlp.linear_fc2.1.00
layers.1.adaLN_mlp_ln.1.00
layers.1.adaLN_sa_ln.1.00
layers.2.1.00
layers.2.self_attention.1.00
layers.2.self_attention.to_q.1.00
layers.2.self_attention.to_k.1.00
layers.2.self_attention.to_v.1.00
layers.2.self_attention.to_out.1.00
layers.2.self_attention.norm_q.1.00
layers.2.self_attention.norm_k.1.00
layers.2.mlp.1.00
layers.2.mlp.gate_proj.1.00
layers.2.mlp.up_proj.1.00
layers.2.mlp.linear_fc2.1.00
layers.2.adaLN_mlp_ln.1.00
layers.2.adaLN_sa_ln.1.00
layers.3.1.00
layers.3.self_attention.1.00
layers.3.self_attention.to_q.1.00
layers.3.self_attention.to_k.1.00
layers.3.self_attention.to_v.1.00
layers.3.self_attention.to_out.1.00
layers.3.self_attention.norm_q.1.00
layers.3.self_attention.norm_k.1.00
layers.3.mlp.1.00
layers.3.mlp.gate_proj.1.00
layers.3.mlp.up_proj.1.00
layers.3.mlp.linear_fc2.1.00
layers.3.adaLN_mlp_ln.1.00
layers.3.adaLN_sa_ln.1.00
layers.4.1.00
layers.4.self_attention.1.00
layers.4.self_attention.to_q.1.00
layers.4.self_attention.to_k.1.00
layers.4.self_attention.to_v.1.00
layers.4.self_attention.to_out.1.00
layers.4.self_attention.norm_q.1.00
layers.4.self_attention.norm_k.1.00
layers.4.mlp.1.00
layers.4.mlp.gate_proj.1.00
layers.4.mlp.up_proj.1.00
layers.4.mlp.linear_fc2.1.00
layers.4.adaLN_mlp_ln.1.00
layers.4.adaLN_sa_ln.1.00
layers.5.1.00
layers.5.self_attention.1.00
layers.5.self_attention.to_q.1.00
layers.5.self_attention.to_k.1.00
layers.5.self_attention.to_v.1.00
layers.5.self_attention.to_out.1.00
layers.5.self_attention.norm_q.1.00
layers.5.self_attention.norm_k.1.00
layers.5.mlp.1.00
layers.5.mlp.gate_proj.1.00
layers.5.mlp.up_proj.1.00
layers.5.mlp.linear_fc2.1.00
layers.5.adaLN_mlp_ln.1.00
layers.5.adaLN_sa_ln.1.00
layers.6.1.00
layers.6.self_attention.1.00
layers.6.self_attention.to_q.1.00
layers.6.self_attention.to_k.1.00
layers.6.self_attention.to_v.1.00
layers.6.self_attention.to_out.1.00
layers.6.self_attention.norm_q.1.00
layers.6.self_attention.norm_k.1.00
layers.6.mlp.1.00
layers.6.mlp.gate_proj.1.00
layers.6.mlp.up_proj.1.00
layers.6.mlp.linear_fc2.1.00
layers.6.adaLN_mlp_ln.1.00
layers.6.adaLN_sa_ln.1.00
layers.7.1.00
layers.7.self_attention.1.00
layers.7.self_attention.to_q.1.00
layers.7.self_attention.to_k.1.00
layers.7.self_attention.to_v.1.00
layers.7.self_attention.to_out.1.00
layers.7.self_attention.norm_q.1.00
layers.7.self_attention.norm_k.1.00
layers.7.mlp.1.00
layers.7.mlp.gate_proj.1.00
layers.7.mlp.up_proj.1.00
layers.7.mlp.linear_fc2.1.00
layers.7.adaLN_mlp_ln.1.00
layers.7.adaLN_sa_ln.1.00
layers.8.1.00
layers.8.self_attention.1.00
layers.8.self_attention.to_q.1.00
layers.8.self_attention.to_k.1.00
layers.8.self_attention.to_v.1.00
layers.8.self_attention.to_out.1.00
layers.8.self_attention.norm_q.1.00
layers.8.self_attention.norm_k.1.00
layers.8.mlp.1.00
layers.8.mlp.gate_proj.1.00
layers.8.mlp.up_proj.1.00
layers.8.mlp.linear_fc2.1.00
layers.8.adaLN_mlp_ln.1.00
layers.8.adaLN_sa_ln.1.00
layers.9.1.00
layers.9.self_attention.1.00
layers.9.self_attention.to_q.1.00
layers.9.self_attention.to_k.1.00
layers.9.self_attention.to_v.1.00
layers.9.self_attention.to_out.1.00
layers.9.self_attention.norm_q.1.00
layers.9.self_attention.norm_k.1.00
layers.9.mlp.1.00
layers.9.mlp.gate_proj.1.00
layers.9.mlp.up_proj.1.00
layers.9.mlp.linear_fc2.1.00
layers.9.adaLN_mlp_ln.1.00
layers.9.adaLN_sa_ln.1.00
layers.10.1.00
layers.10.self_attention.1.00
layers.10.self_attention.to_q.1.00
layers.10.self_attention.to_k.1.00
layers.10.self_attention.to_v.1.00
layers.10.self_attention.to_out.1.00
layers.10.self_attention.norm_q.1.00
layers.10.self_attention.norm_k.1.00
layers.10.mlp.1.00
layers.10.mlp.gate_proj.1.00
layers.10.mlp.up_proj.1.00
layers.10.mlp.linear_fc2.1.00
layers.10.adaLN_mlp_ln.1.00
layers.10.adaLN_sa_ln.1.00
layers.11.1.00
layers.11.self_attention.1.00
layers.11.self_attention.to_q.1.00
layers.11.self_attention.to_k.1.00
layers.11.self_attention.to_v.1.00
layers.11.self_attention.to_out.1.00
layers.11.self_attention.norm_q.1.00
layers.11.self_attention.norm_k.1.00
layers.11.mlp.1.00
layers.11.mlp.gate_proj.1.00
layers.11.mlp.up_proj.1.00
layers.11.mlp.linear_fc2.1.00
layers.11.adaLN_mlp_ln.1.00
layers.11.adaLN_sa_ln.1.00
layers.12.1.00
layers.12.self_attention.1.00
layers.12.self_attention.to_q.1.00
layers.12.self_attention.to_k.1.00
layers.12.self_attention.to_v.1.00
layers.12.self_attention.to_out.1.00
layers.12.self_attention.norm_q.1.00
layers.12.self_attention.norm_k.1.00
layers.12.mlp.1.00
layers.12.mlp.gate_proj.1.00
layers.12.mlp.up_proj.1.00
layers.12.mlp.linear_fc2.1.00
layers.12.adaLN_mlp_ln.1.00
layers.12.adaLN_sa_ln.1.00
layers.13.1.00
layers.13.self_attention.1.00
layers.13.self_attention.to_q.1.00
layers.13.self_attention.to_k.1.00
layers.13.self_attention.to_v.1.00
layers.13.self_attention.to_out.1.00
layers.13.self_attention.norm_q.1.00
layers.13.self_attention.norm_k.1.00
layers.13.mlp.1.00
layers.13.mlp.gate_proj.1.00
layers.13.mlp.up_proj.1.00
layers.13.mlp.linear_fc2.1.00
layers.13.adaLN_mlp_ln.1.00
layers.13.adaLN_sa_ln.1.00
layers.14.1.00
layers.14.self_attention.1.00
layers.14.self_attention.to_q.1.00
layers.14.self_attention.to_k.1.00
layers.14.self_attention.to_v.1.00
layers.14.self_attention.to_out.1.00
layers.14.self_attention.norm_q.1.00
layers.14.self_attention.norm_k.1.00
layers.14.mlp.1.00
layers.14.mlp.gate_proj.1.00
layers.14.mlp.up_proj.1.00
layers.14.mlp.linear_fc2.1.00
layers.14.adaLN_mlp_ln.1.00
layers.14.adaLN_sa_ln.1.00
layers.15.1.00
layers.15.self_attention.1.00
layers.15.self_attention.to_q.1.00
layers.15.self_attention.to_k.1.00
layers.15.self_attention.to_v.1.00
layers.15.self_attention.to_out.1.00
layers.15.self_attention.norm_q.1.00
layers.15.self_attention.norm_k.1.00
layers.15.mlp.1.00
layers.15.mlp.gate_proj.1.00
layers.15.mlp.up_proj.1.00
layers.15.mlp.linear_fc2.1.00
layers.15.adaLN_mlp_ln.1.00
layers.15.adaLN_sa_ln.1.00
layers.16.1.00
layers.16.self_attention.1.00
layers.16.self_attention.to_q.1.00
layers.16.self_attention.to_k.1.00
layers.16.self_attention.to_v.1.00
layers.16.self_attention.to_out.1.00
layers.16.self_attention.norm_q.1.00
layers.16.self_attention.norm_k.1.00
layers.16.mlp.1.00
layers.16.mlp.gate_proj.1.00
layers.16.mlp.up_proj.1.00
layers.16.mlp.linear_fc2.1.00
layers.16.adaLN_mlp_ln.1.00
layers.16.adaLN_sa_ln.1.00
layers.17.1.00
layers.17.self_attention.1.00
layers.17.self_attention.to_q.1.00
layers.17.self_attention.to_k.1.00
layers.17.self_attention.to_v.1.00
layers.17.self_attention.to_out.1.00
layers.17.self_attention.norm_q.1.00
layers.17.self_attention.norm_k.1.00
layers.17.mlp.1.00
layers.17.mlp.gate_proj.1.00
layers.17.mlp.up_proj.1.00
layers.17.mlp.linear_fc2.1.00
layers.17.adaLN_mlp_ln.1.00
layers.17.adaLN_sa_ln.1.00
layers.18.1.00
layers.18.self_attention.1.00
layers.18.self_attention.to_q.1.00
layers.18.self_attention.to_k.1.00
layers.18.self_attention.to_v.1.00
layers.18.self_attention.to_out.1.00
layers.18.self_attention.norm_q.1.00
layers.18.self_attention.norm_k.1.00
layers.18.mlp.1.00
layers.18.mlp.gate_proj.1.00
layers.18.mlp.up_proj.1.00
layers.18.mlp.linear_fc2.1.00
layers.18.adaLN_mlp_ln.1.00
layers.18.adaLN_sa_ln.1.00
layers.19.1.00
layers.19.self_attention.1.00
layers.19.self_attention.to_q.1.00
layers.19.self_attention.to_k.1.00
layers.19.self_attention.to_v.1.00
layers.19.self_attention.to_out.1.00
layers.19.self_attention.norm_q.1.00
layers.19.self_attention.norm_k.1.00
layers.19.mlp.1.00
layers.19.mlp.gate_proj.1.00
layers.19.mlp.up_proj.1.00
layers.19.mlp.linear_fc2.1.00
layers.19.adaLN_mlp_ln.1.00
layers.19.adaLN_sa_ln.1.00
layers.20.1.00
layers.20.self_attention.1.00
layers.20.self_attention.to_q.1.00
layers.20.self_attention.to_k.1.00
layers.20.self_attention.to_v.1.00
layers.20.self_attention.to_out.1.00
layers.20.self_attention.norm_q.1.00
layers.20.self_attention.norm_k.1.00
layers.20.mlp.1.00
layers.20.mlp.gate_proj.1.00
layers.20.mlp.up_proj.1.00
layers.20.mlp.linear_fc2.1.00
layers.20.adaLN_mlp_ln.1.00
layers.20.adaLN_sa_ln.1.00
layers.21.1.00
layers.21.self_attention.1.00
layers.21.self_attention.to_q.1.00
layers.21.self_attention.to_k.1.00
layers.21.self_attention.to_v.1.00
layers.21.self_attention.to_out.1.00
layers.21.self_attention.norm_q.1.00
layers.21.self_attention.norm_k.1.00
layers.21.mlp.1.00
layers.21.mlp.gate_proj.1.00
layers.21.mlp.up_proj.1.00
layers.21.mlp.linear_fc2.1.00
layers.21.adaLN_mlp_ln.1.00
layers.21.adaLN_sa_ln.1.00
layers.22.1.00
layers.22.self_attention.1.00
layers.22.self_attention.to_q.1.00
layers.22.self_attention.to_k.1.00
layers.22.self_attention.to_v.1.00
layers.22.self_attention.to_out.1.00
layers.22.self_attention.norm_q.1.00
layers.22.self_attention.norm_k.1.00
layers.22.mlp.1.00
layers.22.mlp.gate_proj.1.00
layers.22.mlp.up_proj.1.00
layers.22.mlp.linear_fc2.1.00
layers.22.adaLN_mlp_ln.1.00
layers.22.adaLN_sa_ln.1.00
layers.23.1.00
layers.23.self_attention.1.00
layers.23.self_attention.to_q.1.00
layers.23.self_attention.to_k.1.00
layers.23.self_attention.to_v.1.00
layers.23.self_attention.to_out.1.00
layers.23.self_attention.norm_q.1.00
layers.23.self_attention.norm_k.1.00
layers.23.mlp.1.00
layers.23.mlp.gate_proj.1.00
layers.23.mlp.up_proj.1.00
layers.23.mlp.linear_fc2.1.00
layers.23.adaLN_mlp_ln.1.00
layers.23.adaLN_sa_ln.1.00
layers.24.1.00
layers.24.self_attention.1.00
layers.24.self_attention.to_q.1.00
layers.24.self_attention.to_k.1.00
layers.24.self_attention.to_v.1.00
layers.24.self_attention.to_out.1.00
layers.24.self_attention.norm_q.1.00
layers.24.self_attention.norm_k.1.00
layers.24.mlp.1.00
layers.24.mlp.gate_proj.1.00
layers.24.mlp.up_proj.1.00
layers.24.mlp.linear_fc2.1.00
layers.24.adaLN_mlp_ln.1.00
layers.24.adaLN_sa_ln.1.00
layers.25.1.00
layers.25.self_attention.1.00
layers.25.self_attention.to_q.1.00
layers.25.self_attention.to_k.1.00
layers.25.self_attention.to_v.1.00
layers.25.self_attention.to_out.1.00
layers.25.self_attention.norm_q.1.00
layers.25.self_attention.norm_k.1.00
layers.25.mlp.1.00
layers.25.mlp.gate_proj.1.00
layers.25.mlp.up_proj.1.00
layers.25.mlp.linear_fc2.1.00
layers.25.adaLN_mlp_ln.1.00
layers.25.adaLN_sa_ln.1.00
layers.26.1.00
layers.26.self_attention.1.00
layers.26.self_attention.to_q.1.00
layers.26.self_attention.to_k.1.00
layers.26.self_attention.to_v.1.00
layers.26.self_attention.to_out.1.00
layers.26.self_attention.norm_q.1.00
layers.26.self_attention.norm_k.1.00
layers.26.mlp.1.00
layers.26.mlp.gate_proj.1.00
layers.26.mlp.up_proj.1.00
layers.26.mlp.linear_fc2.1.00
layers.26.adaLN_mlp_ln.1.00
layers.26.adaLN_sa_ln.1.00
layers.27.1.00
layers.27.self_attention.1.00
layers.27.self_attention.to_q.1.00
layers.27.self_attention.to_k.1.00
layers.27.self_attention.to_v.1.00
layers.27.self_attention.to_out.1.00
layers.27.self_attention.norm_q.1.00
layers.27.self_attention.norm_k.1.00
layers.27.mlp.1.00
layers.27.mlp.gate_proj.1.00
layers.27.mlp.up_proj.1.00
layers.27.mlp.linear_fc2.1.00
layers.27.adaLN_mlp_ln.1.00
layers.27.adaLN_sa_ln.1.00
layers.28.1.00
layers.28.self_attention.1.00
layers.28.self_attention.to_q.1.00
layers.28.self_attention.to_k.1.00
layers.28.self_attention.to_v.1.00
layers.28.self_attention.to_out.1.00
layers.28.self_attention.norm_q.1.00
layers.28.self_attention.norm_k.1.00
layers.28.mlp.1.00
layers.28.mlp.gate_proj.1.00
layers.28.mlp.up_proj.1.00
layers.28.mlp.linear_fc2.1.00
layers.28.adaLN_mlp_ln.1.00
layers.28.adaLN_sa_ln.1.00
layers.29.1.00
layers.29.self_attention.1.00
layers.29.self_attention.to_q.1.00
layers.29.self_attention.to_k.1.00
layers.29.self_attention.to_v.1.00
layers.29.self_attention.to_out.1.00
layers.29.self_attention.norm_q.1.00
layers.29.self_attention.norm_k.1.00
layers.29.mlp.1.00
layers.29.mlp.gate_proj.1.00
layers.29.mlp.up_proj.1.00
layers.29.mlp.linear_fc2.1.00
layers.29.adaLN_mlp_ln.1.00
layers.29.adaLN_sa_ln.1.00
layers.30.1.00
layers.30.self_attention.1.00
layers.30.self_attention.to_q.1.00
layers.30.self_attention.to_k.1.00
layers.30.self_attention.to_v.1.00
layers.30.self_attention.to_out.1.00
layers.30.self_attention.norm_q.1.00
layers.30.self_attention.norm_k.1.00
layers.30.mlp.1.00
layers.30.mlp.gate_proj.1.00
layers.30.mlp.up_proj.1.00
layers.30.mlp.linear_fc2.1.00
layers.30.adaLN_mlp_ln.1.00
layers.30.adaLN_sa_ln.1.00
layers.31.1.00
layers.31.self_attention.1.00
layers.31.self_attention.to_q.1.00
layers.31.self_attention.to_k.1.00
layers.31.self_attention.to_v.1.00
layers.31.self_attention.to_out.1.00
layers.31.self_attention.norm_q.1.00
layers.31.self_attention.norm_k.1.00
layers.31.mlp.1.00
layers.31.mlp.gate_proj.1.00
layers.31.mlp.up_proj.1.00
layers.31.mlp.linear_fc2.1.00
layers.31.adaLN_mlp_ln.1.00
layers.31.adaLN_sa_ln.1.00
layers.32.1.00
layers.32.self_attention.1.00
layers.32.self_attention.to_q.1.00
layers.32.self_attention.to_k.1.00
layers.32.self_attention.to_v.1.00
layers.32.self_attention.to_out.1.00
layers.32.self_attention.norm_q.1.00
layers.32.self_attention.norm_k.1.00
layers.32.mlp.1.00
layers.32.mlp.gate_proj.1.00
layers.32.mlp.up_proj.1.00
layers.32.mlp.linear_fc2.1.00
layers.32.adaLN_mlp_ln.1.00
layers.32.adaLN_sa_ln.1.00
layers.33.1.00
layers.33.self_attention.1.00
layers.33.self_attention.to_q.1.00
layers.33.self_attention.to_k.1.00
layers.33.self_attention.to_v.1.00
layers.33.self_attention.to_out.1.00
layers.33.self_attention.norm_q.1.00
layers.33.self_attention.norm_k.1.00
layers.33.mlp.1.00
layers.33.mlp.gate_proj.1.00
layers.33.mlp.up_proj.1.00
layers.33.mlp.linear_fc2.1.00
layers.33.adaLN_mlp_ln.1.00
layers.33.adaLN_sa_ln.1.00
layers.34.1.00
layers.34.self_attention.1.00
layers.34.self_attention.to_q.1.00
layers.34.self_attention.to_k.1.00
layers.34.self_attention.to_v.1.00
layers.34.self_attention.to_out.1.00
layers.34.self_attention.norm_q.1.00
layers.34.self_attention.norm_k.1.00
layers.34.mlp.1.00
layers.34.mlp.gate_proj.1.00
layers.34.mlp.up_proj.1.00
layers.34.mlp.linear_fc2.1.00
layers.34.adaLN_mlp_ln.1.00
layers.34.adaLN_sa_ln.1.00
layers.35.1.00
layers.35.self_attention.1.00
layers.35.self_attention.to_q.1.00
layers.35.self_attention.to_k.1.00
layers.35.self_attention.to_v.1.00
layers.35.self_attention.to_out.1.00
layers.35.self_attention.norm_q.1.00
layers.35.self_attention.norm_k.1.00
layers.35.mlp.1.00
layers.35.mlp.gate_proj.1.00
layers.35.mlp.up_proj.1.00
layers.35.mlp.linear_fc2.1.00
layers.35.adaLN_mlp_ln.1.00
layers.35.adaLN_sa_ln.1.00
adaLN_modulation.1.00
final_norm.1.00
final_linear.1.00

ERNIE-Image is Baidu's 8B Apache-2.0 model - the one that genuinely impressed people with structured layout and text-in-image before the community's attention moved on. If you're one of the people still using it, Model Scale ERNIE Image is the fine-grained scaler for it: one slider per layer and per attention component, covering all 36 transformer blocks of ernie-image.safetensors. This is the most granular scaler in the pack, and it exists because ERNIE-Image's strength - following structured prompts - isn't spread evenly across its weights.

The mechanism is the same as every scale node here: clone the model, match each weight to the longest matching widget prefix, and scale by weight × scale via add_patches. 1.0 leaves a layer untouched, 0.0 zeroes it out, values above 1.0 amplify. All sliders default to 1.0, so the node is inert until you move something. Input is a model, output is a single scaled MODEL for your sampler.

The inputs, grouped

The schema is huge - hundreds of floats - because every layer is expanded into its sub-components. But they organize cleanly:

  • layers.N - the per-block master slider. Start here; the sub-sliders are for when you need finer control.
  • layers.N.self_attention.to_q / .to_k / .to_v / .to_out - the individual attention projection matrices inside each block. This is the rare level of control: you can weaken the query projections across all blocks without touching the MLPs, which is a very different surgical move than scaling whole blocks.
  • layers.N.mlp.gate_proj / .up_proj / .linear_fc2 - the feed-forward path, the part most responsible for factual/layout recall in these architectures.
  • layers.N.adaLN_mlp_ln / .adaLN_sa_ln - the adaptive normalization layers that modulate each block.
  • Top-level: x_embedder, text_proj, time_embedding, adaLN_modulation, final_norm, final_linear.

What to actually try

If you're chasing ERNIE-Image's text-rendering quality, the community's finding is that its competence is real but front-loaded - the mid-to-late blocks and the MLP path carry a lot of it. A reasonable first experiment: amplify mlp sub-components in the middle third (layers ~12–24) slightly above 1.0 and see if layout fidelity sharpens, then dial back to taste. Zeroing whole late blocks usually just degrades output. And if the grid-pattern artifact that people noted with ERNIE-Image is bothering you, experimenting with the attention projections in the early layers is a more targeted lever than a global scale.

Install

Search Easygoing in the ComfyUI Manager, or:

cd ComfyUI/custom_nodes
git clone https://github.com/easygoing0114/ComfyUI-easygoing-nodes.git

Restart ComfyUI. No pip extras; needs a current ComfyUI with the V3 node API for the pack to register.

Honest warning

This is a scaler for people already invested in ERNIE-Image, and the sub-component granularity is genuinely useful there - but the node only makes sense if the widget prefixes match your model's actual keys. Run it once with the pack's Key Name Inspector wired in parallel if you're not sure; a mismatch shows up as sliders that do nothing, and that's a wasted afternoon otherwise. Scaling is also in-graph only - save the workflow, or route through the pack's save-with-original nodes if you want a file.

Categoryadvanced/model_merging/model_specific

Inputs (511)

NameTypeDefaultDescription
modelMODEL
x_embedder.FLOAT1.000–2
text_proj.FLOAT1.000–2
time_embedding.FLOAT1.000–2
layers.0.FLOAT1.000–2
layers.0.self_attention.FLOAT1.000–2
layers.0.self_attention.to_q.FLOAT1.000–2
layers.0.self_attention.to_k.FLOAT1.000–2
layers.0.self_attention.to_v.FLOAT1.000–2
layers.0.self_attention.to_out.FLOAT1.000–2
layers.0.self_attention.norm_q.FLOAT1.000–2
layers.0.self_attention.norm_k.FLOAT1.000–2
layers.0.mlp.FLOAT1.000–2
layers.0.mlp.gate_proj.FLOAT1.000–2
layers.0.mlp.up_proj.FLOAT1.000–2
layers.0.mlp.linear_fc2.FLOAT1.000–2
layers.0.adaLN_mlp_ln.FLOAT1.000–2
layers.0.adaLN_sa_ln.FLOAT1.000–2
layers.1.FLOAT1.000–2
layers.1.self_attention.FLOAT1.000–2
layers.1.self_attention.to_q.FLOAT1.000–2
layers.1.self_attention.to_k.FLOAT1.000–2
layers.1.self_attention.to_v.FLOAT1.000–2
layers.1.self_attention.to_out.FLOAT1.000–2
layers.1.self_attention.norm_q.FLOAT1.000–2
layers.1.self_attention.norm_k.FLOAT1.000–2
layers.1.mlp.FLOAT1.000–2
layers.1.mlp.gate_proj.FLOAT1.000–2
layers.1.mlp.up_proj.FLOAT1.000–2
layers.1.mlp.linear_fc2.FLOAT1.000–2
layers.1.adaLN_mlp_ln.FLOAT1.000–2
layers.1.adaLN_sa_ln.FLOAT1.000–2
layers.2.FLOAT1.000–2
layers.2.self_attention.FLOAT1.000–2
layers.2.self_attention.to_q.FLOAT1.000–2
layers.2.self_attention.to_k.FLOAT1.000–2
layers.2.self_attention.to_v.FLOAT1.000–2
layers.2.self_attention.to_out.FLOAT1.000–2
layers.2.self_attention.norm_q.FLOAT1.000–2
layers.2.self_attention.norm_k.FLOAT1.000–2
layers.2.mlp.FLOAT1.000–2
layers.2.mlp.gate_proj.FLOAT1.000–2
layers.2.mlp.up_proj.FLOAT1.000–2
layers.2.mlp.linear_fc2.FLOAT1.000–2
layers.2.adaLN_mlp_ln.FLOAT1.000–2
layers.2.adaLN_sa_ln.FLOAT1.000–2
layers.3.FLOAT1.000–2
layers.3.self_attention.FLOAT1.000–2
layers.3.self_attention.to_q.FLOAT1.000–2
layers.3.self_attention.to_k.FLOAT1.000–2
layers.3.self_attention.to_v.FLOAT1.000–2
layers.3.self_attention.to_out.FLOAT1.000–2
layers.3.self_attention.norm_q.FLOAT1.000–2
layers.3.self_attention.norm_k.FLOAT1.000–2
layers.3.mlp.FLOAT1.000–2
layers.3.mlp.gate_proj.FLOAT1.000–2
layers.3.mlp.up_proj.FLOAT1.000–2
layers.3.mlp.linear_fc2.FLOAT1.000–2
layers.3.adaLN_mlp_ln.FLOAT1.000–2
layers.3.adaLN_sa_ln.FLOAT1.000–2
layers.4.FLOAT1.000–2
layers.4.self_attention.FLOAT1.000–2
layers.4.self_attention.to_q.FLOAT1.000–2
layers.4.self_attention.to_k.FLOAT1.000–2
layers.4.self_attention.to_v.FLOAT1.000–2
layers.4.self_attention.to_out.FLOAT1.000–2
layers.4.self_attention.norm_q.FLOAT1.000–2
layers.4.self_attention.norm_k.FLOAT1.000–2
layers.4.mlp.FLOAT1.000–2
layers.4.mlp.gate_proj.FLOAT1.000–2
layers.4.mlp.up_proj.FLOAT1.000–2
layers.4.mlp.linear_fc2.FLOAT1.000–2
layers.4.adaLN_mlp_ln.FLOAT1.000–2
layers.4.adaLN_sa_ln.FLOAT1.000–2
layers.5.FLOAT1.000–2
layers.5.self_attention.FLOAT1.000–2
layers.5.self_attention.to_q.FLOAT1.000–2
layers.5.self_attention.to_k.FLOAT1.000–2
layers.5.self_attention.to_v.FLOAT1.000–2
layers.5.self_attention.to_out.FLOAT1.000–2
layers.5.self_attention.norm_q.FLOAT1.000–2
layers.5.self_attention.norm_k.FLOAT1.000–2
layers.5.mlp.FLOAT1.000–2
layers.5.mlp.gate_proj.FLOAT1.000–2
layers.5.mlp.up_proj.FLOAT1.000–2
layers.5.mlp.linear_fc2.FLOAT1.000–2
layers.5.adaLN_mlp_ln.FLOAT1.000–2
layers.5.adaLN_sa_ln.FLOAT1.000–2
layers.6.FLOAT1.000–2
layers.6.self_attention.FLOAT1.000–2
layers.6.self_attention.to_q.FLOAT1.000–2
layers.6.self_attention.to_k.FLOAT1.000–2
layers.6.self_attention.to_v.FLOAT1.000–2
layers.6.self_attention.to_out.FLOAT1.000–2
layers.6.self_attention.norm_q.FLOAT1.000–2
layers.6.self_attention.norm_k.FLOAT1.000–2
layers.6.mlp.FLOAT1.000–2
layers.6.mlp.gate_proj.FLOAT1.000–2
layers.6.mlp.up_proj.FLOAT1.000–2
layers.6.mlp.linear_fc2.FLOAT1.000–2
layers.6.adaLN_mlp_ln.FLOAT1.000–2
layers.6.adaLN_sa_ln.FLOAT1.000–2
layers.7.FLOAT1.000–2
layers.7.self_attention.FLOAT1.000–2
layers.7.self_attention.to_q.FLOAT1.000–2
layers.7.self_attention.to_k.FLOAT1.000–2
layers.7.self_attention.to_v.FLOAT1.000–2
layers.7.self_attention.to_out.FLOAT1.000–2
layers.7.self_attention.norm_q.FLOAT1.000–2
layers.7.self_attention.norm_k.FLOAT1.000–2
layers.7.mlp.FLOAT1.000–2
layers.7.mlp.gate_proj.FLOAT1.000–2
layers.7.mlp.up_proj.FLOAT1.000–2
layers.7.mlp.linear_fc2.FLOAT1.000–2
layers.7.adaLN_mlp_ln.FLOAT1.000–2
layers.7.adaLN_sa_ln.FLOAT1.000–2
layers.8.FLOAT1.000–2
layers.8.self_attention.FLOAT1.000–2
layers.8.self_attention.to_q.FLOAT1.000–2
layers.8.self_attention.to_k.FLOAT1.000–2
layers.8.self_attention.to_v.FLOAT1.000–2
layers.8.self_attention.to_out.FLOAT1.000–2
layers.8.self_attention.norm_q.FLOAT1.000–2
layers.8.self_attention.norm_k.FLOAT1.000–2
layers.8.mlp.FLOAT1.000–2
layers.8.mlp.gate_proj.FLOAT1.000–2
layers.8.mlp.up_proj.FLOAT1.000–2
layers.8.mlp.linear_fc2.FLOAT1.000–2
layers.8.adaLN_mlp_ln.FLOAT1.000–2
layers.8.adaLN_sa_ln.FLOAT1.000–2
layers.9.FLOAT1.000–2
layers.9.self_attention.FLOAT1.000–2
layers.9.self_attention.to_q.FLOAT1.000–2
layers.9.self_attention.to_k.FLOAT1.000–2
layers.9.self_attention.to_v.FLOAT1.000–2
layers.9.self_attention.to_out.FLOAT1.000–2
layers.9.self_attention.norm_q.FLOAT1.000–2
layers.9.self_attention.norm_k.FLOAT1.000–2
layers.9.mlp.FLOAT1.000–2
layers.9.mlp.gate_proj.FLOAT1.000–2
layers.9.mlp.up_proj.FLOAT1.000–2
layers.9.mlp.linear_fc2.FLOAT1.000–2
layers.9.adaLN_mlp_ln.FLOAT1.000–2
layers.9.adaLN_sa_ln.FLOAT1.000–2
layers.10.FLOAT1.000–2
layers.10.self_attention.FLOAT1.000–2
layers.10.self_attention.to_q.FLOAT1.000–2
layers.10.self_attention.to_k.FLOAT1.000–2
layers.10.self_attention.to_v.FLOAT1.000–2
layers.10.self_attention.to_out.FLOAT1.000–2
layers.10.self_attention.norm_q.FLOAT1.000–2
layers.10.self_attention.norm_k.FLOAT1.000–2
layers.10.mlp.FLOAT1.000–2
layers.10.mlp.gate_proj.FLOAT1.000–2
layers.10.mlp.up_proj.FLOAT1.000–2
layers.10.mlp.linear_fc2.FLOAT1.000–2
layers.10.adaLN_mlp_ln.FLOAT1.000–2
layers.10.adaLN_sa_ln.FLOAT1.000–2
layers.11.FLOAT1.000–2
layers.11.self_attention.FLOAT1.000–2
layers.11.self_attention.to_q.FLOAT1.000–2
layers.11.self_attention.to_k.FLOAT1.000–2
layers.11.self_attention.to_v.FLOAT1.000–2
layers.11.self_attention.to_out.FLOAT1.000–2
layers.11.self_attention.norm_q.FLOAT1.000–2
layers.11.self_attention.norm_k.FLOAT1.000–2
layers.11.mlp.FLOAT1.000–2
layers.11.mlp.gate_proj.FLOAT1.000–2
layers.11.mlp.up_proj.FLOAT1.000–2
layers.11.mlp.linear_fc2.FLOAT1.000–2
layers.11.adaLN_mlp_ln.FLOAT1.000–2
layers.11.adaLN_sa_ln.FLOAT1.000–2
layers.12.FLOAT1.000–2
layers.12.self_attention.FLOAT1.000–2
layers.12.self_attention.to_q.FLOAT1.000–2
layers.12.self_attention.to_k.FLOAT1.000–2
layers.12.self_attention.to_v.FLOAT1.000–2
layers.12.self_attention.to_out.FLOAT1.000–2
layers.12.self_attention.norm_q.FLOAT1.000–2
layers.12.self_attention.norm_k.FLOAT1.000–2
layers.12.mlp.FLOAT1.000–2
layers.12.mlp.gate_proj.FLOAT1.000–2
layers.12.mlp.up_proj.FLOAT1.000–2
layers.12.mlp.linear_fc2.FLOAT1.000–2
layers.12.adaLN_mlp_ln.FLOAT1.000–2
layers.12.adaLN_sa_ln.FLOAT1.000–2
layers.13.FLOAT1.000–2
layers.13.self_attention.FLOAT1.000–2
layers.13.self_attention.to_q.FLOAT1.000–2
layers.13.self_attention.to_k.FLOAT1.000–2
layers.13.self_attention.to_v.FLOAT1.000–2
layers.13.self_attention.to_out.FLOAT1.000–2
layers.13.self_attention.norm_q.FLOAT1.000–2
layers.13.self_attention.norm_k.FLOAT1.000–2
layers.13.mlp.FLOAT1.000–2
layers.13.mlp.gate_proj.FLOAT1.000–2
layers.13.mlp.up_proj.FLOAT1.000–2
layers.13.mlp.linear_fc2.FLOAT1.000–2
layers.13.adaLN_mlp_ln.FLOAT1.000–2
layers.13.adaLN_sa_ln.FLOAT1.000–2
layers.14.FLOAT1.000–2
layers.14.self_attention.FLOAT1.000–2
layers.14.self_attention.to_q.FLOAT1.000–2
layers.14.self_attention.to_k.FLOAT1.000–2
layers.14.self_attention.to_v.FLOAT1.000–2
layers.14.self_attention.to_out.FLOAT1.000–2
layers.14.self_attention.norm_q.FLOAT1.000–2
layers.14.self_attention.norm_k.FLOAT1.000–2
layers.14.mlp.FLOAT1.000–2
layers.14.mlp.gate_proj.FLOAT1.000–2
layers.14.mlp.up_proj.FLOAT1.000–2
layers.14.mlp.linear_fc2.FLOAT1.000–2
layers.14.adaLN_mlp_ln.FLOAT1.000–2
layers.14.adaLN_sa_ln.FLOAT1.000–2
layers.15.FLOAT1.000–2
layers.15.self_attention.FLOAT1.000–2
layers.15.self_attention.to_q.FLOAT1.000–2
layers.15.self_attention.to_k.FLOAT1.000–2
layers.15.self_attention.to_v.FLOAT1.000–2
layers.15.self_attention.to_out.FLOAT1.000–2
layers.15.self_attention.norm_q.FLOAT1.000–2
layers.15.self_attention.norm_k.FLOAT1.000–2
layers.15.mlp.FLOAT1.000–2
layers.15.mlp.gate_proj.FLOAT1.000–2
layers.15.mlp.up_proj.FLOAT1.000–2
layers.15.mlp.linear_fc2.FLOAT1.000–2
layers.15.adaLN_mlp_ln.FLOAT1.000–2
layers.15.adaLN_sa_ln.FLOAT1.000–2
layers.16.FLOAT1.000–2
layers.16.self_attention.FLOAT1.000–2
layers.16.self_attention.to_q.FLOAT1.000–2
layers.16.self_attention.to_k.FLOAT1.000–2
layers.16.self_attention.to_v.FLOAT1.000–2
layers.16.self_attention.to_out.FLOAT1.000–2
layers.16.self_attention.norm_q.FLOAT1.000–2
layers.16.self_attention.norm_k.FLOAT1.000–2
layers.16.mlp.FLOAT1.000–2
layers.16.mlp.gate_proj.FLOAT1.000–2
layers.16.mlp.up_proj.FLOAT1.000–2
layers.16.mlp.linear_fc2.FLOAT1.000–2
layers.16.adaLN_mlp_ln.FLOAT1.000–2
layers.16.adaLN_sa_ln.FLOAT1.000–2
layers.17.FLOAT1.000–2
layers.17.self_attention.FLOAT1.000–2
layers.17.self_attention.to_q.FLOAT1.000–2
layers.17.self_attention.to_k.FLOAT1.000–2
layers.17.self_attention.to_v.FLOAT1.000–2
layers.17.self_attention.to_out.FLOAT1.000–2
layers.17.self_attention.norm_q.FLOAT1.000–2
layers.17.self_attention.norm_k.FLOAT1.000–2
layers.17.mlp.FLOAT1.000–2
layers.17.mlp.gate_proj.FLOAT1.000–2
layers.17.mlp.up_proj.FLOAT1.000–2
layers.17.mlp.linear_fc2.FLOAT1.000–2
layers.17.adaLN_mlp_ln.FLOAT1.000–2
layers.17.adaLN_sa_ln.FLOAT1.000–2
layers.18.FLOAT1.000–2
layers.18.self_attention.FLOAT1.000–2
layers.18.self_attention.to_q.FLOAT1.000–2
layers.18.self_attention.to_k.FLOAT1.000–2
layers.18.self_attention.to_v.FLOAT1.000–2
layers.18.self_attention.to_out.FLOAT1.000–2
layers.18.self_attention.norm_q.FLOAT1.000–2
layers.18.self_attention.norm_k.FLOAT1.000–2
layers.18.mlp.FLOAT1.000–2
layers.18.mlp.gate_proj.FLOAT1.000–2
layers.18.mlp.up_proj.FLOAT1.000–2
layers.18.mlp.linear_fc2.FLOAT1.000–2
layers.18.adaLN_mlp_ln.FLOAT1.000–2
layers.18.adaLN_sa_ln.FLOAT1.000–2
layers.19.FLOAT1.000–2
layers.19.self_attention.FLOAT1.000–2
layers.19.self_attention.to_q.FLOAT1.000–2
layers.19.self_attention.to_k.FLOAT1.000–2
layers.19.self_attention.to_v.FLOAT1.000–2
layers.19.self_attention.to_out.FLOAT1.000–2
layers.19.self_attention.norm_q.FLOAT1.000–2
layers.19.self_attention.norm_k.FLOAT1.000–2
layers.19.mlp.FLOAT1.000–2
layers.19.mlp.gate_proj.FLOAT1.000–2
layers.19.mlp.up_proj.FLOAT1.000–2
layers.19.mlp.linear_fc2.FLOAT1.000–2
layers.19.adaLN_mlp_ln.FLOAT1.000–2
layers.19.adaLN_sa_ln.FLOAT1.000–2
layers.20.FLOAT1.000–2
layers.20.self_attention.FLOAT1.000–2
layers.20.self_attention.to_q.FLOAT1.000–2
layers.20.self_attention.to_k.FLOAT1.000–2
layers.20.self_attention.to_v.FLOAT1.000–2
layers.20.self_attention.to_out.FLOAT1.000–2
layers.20.self_attention.norm_q.FLOAT1.000–2
layers.20.self_attention.norm_k.FLOAT1.000–2
layers.20.mlp.FLOAT1.000–2
layers.20.mlp.gate_proj.FLOAT1.000–2
layers.20.mlp.up_proj.FLOAT1.000–2
layers.20.mlp.linear_fc2.FLOAT1.000–2
layers.20.adaLN_mlp_ln.FLOAT1.000–2
layers.20.adaLN_sa_ln.FLOAT1.000–2
layers.21.FLOAT1.000–2
layers.21.self_attention.FLOAT1.000–2
layers.21.self_attention.to_q.FLOAT1.000–2
layers.21.self_attention.to_k.FLOAT1.000–2
layers.21.self_attention.to_v.FLOAT1.000–2
layers.21.self_attention.to_out.FLOAT1.000–2
layers.21.self_attention.norm_q.FLOAT1.000–2
layers.21.self_attention.norm_k.FLOAT1.000–2
layers.21.mlp.FLOAT1.000–2
layers.21.mlp.gate_proj.FLOAT1.000–2
layers.21.mlp.up_proj.FLOAT1.000–2
layers.21.mlp.linear_fc2.FLOAT1.000–2
layers.21.adaLN_mlp_ln.FLOAT1.000–2
layers.21.adaLN_sa_ln.FLOAT1.000–2
layers.22.FLOAT1.000–2
layers.22.self_attention.FLOAT1.000–2
layers.22.self_attention.to_q.FLOAT1.000–2
layers.22.self_attention.to_k.FLOAT1.000–2
layers.22.self_attention.to_v.FLOAT1.000–2
layers.22.self_attention.to_out.FLOAT1.000–2
layers.22.self_attention.norm_q.FLOAT1.000–2
layers.22.self_attention.norm_k.FLOAT1.000–2
layers.22.mlp.FLOAT1.000–2
layers.22.mlp.gate_proj.FLOAT1.000–2
layers.22.mlp.up_proj.FLOAT1.000–2
layers.22.mlp.linear_fc2.FLOAT1.000–2
layers.22.adaLN_mlp_ln.FLOAT1.000–2
layers.22.adaLN_sa_ln.FLOAT1.000–2
layers.23.FLOAT1.000–2
layers.23.self_attention.FLOAT1.000–2
layers.23.self_attention.to_q.FLOAT1.000–2
layers.23.self_attention.to_k.FLOAT1.000–2
layers.23.self_attention.to_v.FLOAT1.000–2
layers.23.self_attention.to_out.FLOAT1.000–2
layers.23.self_attention.norm_q.FLOAT1.000–2
layers.23.self_attention.norm_k.FLOAT1.000–2
layers.23.mlp.FLOAT1.000–2
layers.23.mlp.gate_proj.FLOAT1.000–2
layers.23.mlp.up_proj.FLOAT1.000–2
layers.23.mlp.linear_fc2.FLOAT1.000–2
layers.23.adaLN_mlp_ln.FLOAT1.000–2
layers.23.adaLN_sa_ln.FLOAT1.000–2
layers.24.FLOAT1.000–2
layers.24.self_attention.FLOAT1.000–2
layers.24.self_attention.to_q.FLOAT1.000–2
layers.24.self_attention.to_k.FLOAT1.000–2
layers.24.self_attention.to_v.FLOAT1.000–2
layers.24.self_attention.to_out.FLOAT1.000–2
layers.24.self_attention.norm_q.FLOAT1.000–2
layers.24.self_attention.norm_k.FLOAT1.000–2
layers.24.mlp.FLOAT1.000–2
layers.24.mlp.gate_proj.FLOAT1.000–2
layers.24.mlp.up_proj.FLOAT1.000–2
layers.24.mlp.linear_fc2.FLOAT1.000–2
layers.24.adaLN_mlp_ln.FLOAT1.000–2
layers.24.adaLN_sa_ln.FLOAT1.000–2
layers.25.FLOAT1.000–2
layers.25.self_attention.FLOAT1.000–2
layers.25.self_attention.to_q.FLOAT1.000–2
layers.25.self_attention.to_k.FLOAT1.000–2
layers.25.self_attention.to_v.FLOAT1.000–2
layers.25.self_attention.to_out.FLOAT1.000–2
layers.25.self_attention.norm_q.FLOAT1.000–2
layers.25.self_attention.norm_k.FLOAT1.000–2
layers.25.mlp.FLOAT1.000–2
layers.25.mlp.gate_proj.FLOAT1.000–2
layers.25.mlp.up_proj.FLOAT1.000–2
layers.25.mlp.linear_fc2.FLOAT1.000–2
layers.25.adaLN_mlp_ln.FLOAT1.000–2
layers.25.adaLN_sa_ln.FLOAT1.000–2
layers.26.FLOAT1.000–2
layers.26.self_attention.FLOAT1.000–2
layers.26.self_attention.to_q.FLOAT1.000–2
layers.26.self_attention.to_k.FLOAT1.000–2
layers.26.self_attention.to_v.FLOAT1.000–2
layers.26.self_attention.to_out.FLOAT1.000–2
layers.26.self_attention.norm_q.FLOAT1.000–2
layers.26.self_attention.norm_k.FLOAT1.000–2
layers.26.mlp.FLOAT1.000–2
layers.26.mlp.gate_proj.FLOAT1.000–2
layers.26.mlp.up_proj.FLOAT1.000–2
layers.26.mlp.linear_fc2.FLOAT1.000–2
layers.26.adaLN_mlp_ln.FLOAT1.000–2
layers.26.adaLN_sa_ln.FLOAT1.000–2
layers.27.FLOAT1.000–2
layers.27.self_attention.FLOAT1.000–2
layers.27.self_attention.to_q.FLOAT1.000–2
layers.27.self_attention.to_k.FLOAT1.000–2
layers.27.self_attention.to_v.FLOAT1.000–2
layers.27.self_attention.to_out.FLOAT1.000–2
layers.27.self_attention.norm_q.FLOAT1.000–2
layers.27.self_attention.norm_k.FLOAT1.000–2
layers.27.mlp.FLOAT1.000–2
layers.27.mlp.gate_proj.FLOAT1.000–2
layers.27.mlp.up_proj.FLOAT1.000–2
layers.27.mlp.linear_fc2.FLOAT1.000–2
layers.27.adaLN_mlp_ln.FLOAT1.000–2
layers.27.adaLN_sa_ln.FLOAT1.000–2
layers.28.FLOAT1.000–2
layers.28.self_attention.FLOAT1.000–2
layers.28.self_attention.to_q.FLOAT1.000–2
layers.28.self_attention.to_k.FLOAT1.000–2
layers.28.self_attention.to_v.FLOAT1.000–2
layers.28.self_attention.to_out.FLOAT1.000–2
layers.28.self_attention.norm_q.FLOAT1.000–2
layers.28.self_attention.norm_k.FLOAT1.000–2
layers.28.mlp.FLOAT1.000–2
layers.28.mlp.gate_proj.FLOAT1.000–2
layers.28.mlp.up_proj.FLOAT1.000–2
layers.28.mlp.linear_fc2.FLOAT1.000–2
layers.28.adaLN_mlp_ln.FLOAT1.000–2
layers.28.adaLN_sa_ln.FLOAT1.000–2
layers.29.FLOAT1.000–2
layers.29.self_attention.FLOAT1.000–2
layers.29.self_attention.to_q.FLOAT1.000–2
layers.29.self_attention.to_k.FLOAT1.000–2
layers.29.self_attention.to_v.FLOAT1.000–2
layers.29.self_attention.to_out.FLOAT1.000–2
layers.29.self_attention.norm_q.FLOAT1.000–2
layers.29.self_attention.norm_k.FLOAT1.000–2
layers.29.mlp.FLOAT1.000–2
layers.29.mlp.gate_proj.FLOAT1.000–2
layers.29.mlp.up_proj.FLOAT1.000–2
layers.29.mlp.linear_fc2.FLOAT1.000–2
layers.29.adaLN_mlp_ln.FLOAT1.000–2
layers.29.adaLN_sa_ln.FLOAT1.000–2
layers.30.FLOAT1.000–2
layers.30.self_attention.FLOAT1.000–2
layers.30.self_attention.to_q.FLOAT1.000–2
layers.30.self_attention.to_k.FLOAT1.000–2
layers.30.self_attention.to_v.FLOAT1.000–2
layers.30.self_attention.to_out.FLOAT1.000–2
layers.30.self_attention.norm_q.FLOAT1.000–2
layers.30.self_attention.norm_k.FLOAT1.000–2
layers.30.mlp.FLOAT1.000–2
layers.30.mlp.gate_proj.FLOAT1.000–2
layers.30.mlp.up_proj.FLOAT1.000–2
layers.30.mlp.linear_fc2.FLOAT1.000–2
layers.30.adaLN_mlp_ln.FLOAT1.000–2
layers.30.adaLN_sa_ln.FLOAT1.000–2
layers.31.FLOAT1.000–2
layers.31.self_attention.FLOAT1.000–2
layers.31.self_attention.to_q.FLOAT1.000–2
layers.31.self_attention.to_k.FLOAT1.000–2
layers.31.self_attention.to_v.FLOAT1.000–2
layers.31.self_attention.to_out.FLOAT1.000–2
layers.31.self_attention.norm_q.FLOAT1.000–2
layers.31.self_attention.norm_k.FLOAT1.000–2
layers.31.mlp.FLOAT1.000–2
layers.31.mlp.gate_proj.FLOAT1.000–2
layers.31.mlp.up_proj.FLOAT1.000–2
layers.31.mlp.linear_fc2.FLOAT1.000–2
layers.31.adaLN_mlp_ln.FLOAT1.000–2
layers.31.adaLN_sa_ln.FLOAT1.000–2
layers.32.FLOAT1.000–2
layers.32.self_attention.FLOAT1.000–2
layers.32.self_attention.to_q.FLOAT1.000–2
layers.32.self_attention.to_k.FLOAT1.000–2
layers.32.self_attention.to_v.FLOAT1.000–2
layers.32.self_attention.to_out.FLOAT1.000–2
layers.32.self_attention.norm_q.FLOAT1.000–2
layers.32.self_attention.norm_k.FLOAT1.000–2
layers.32.mlp.FLOAT1.000–2
layers.32.mlp.gate_proj.FLOAT1.000–2
layers.32.mlp.up_proj.FLOAT1.000–2
layers.32.mlp.linear_fc2.FLOAT1.000–2
layers.32.adaLN_mlp_ln.FLOAT1.000–2
layers.32.adaLN_sa_ln.FLOAT1.000–2
layers.33.FLOAT1.000–2
layers.33.self_attention.FLOAT1.000–2
layers.33.self_attention.to_q.FLOAT1.000–2
layers.33.self_attention.to_k.FLOAT1.000–2
layers.33.self_attention.to_v.FLOAT1.000–2
layers.33.self_attention.to_out.FLOAT1.000–2
layers.33.self_attention.norm_q.FLOAT1.000–2
layers.33.self_attention.norm_k.FLOAT1.000–2
layers.33.mlp.FLOAT1.000–2
layers.33.mlp.gate_proj.FLOAT1.000–2
layers.33.mlp.up_proj.FLOAT1.000–2
layers.33.mlp.linear_fc2.FLOAT1.000–2
layers.33.adaLN_mlp_ln.FLOAT1.000–2
layers.33.adaLN_sa_ln.FLOAT1.000–2
layers.34.FLOAT1.000–2
layers.34.self_attention.FLOAT1.000–2
layers.34.self_attention.to_q.FLOAT1.000–2
layers.34.self_attention.to_k.FLOAT1.000–2
layers.34.self_attention.to_v.FLOAT1.000–2
layers.34.self_attention.to_out.FLOAT1.000–2
layers.34.self_attention.norm_q.FLOAT1.000–2
layers.34.self_attention.norm_k.FLOAT1.000–2
layers.34.mlp.FLOAT1.000–2
layers.34.mlp.gate_proj.FLOAT1.000–2
layers.34.mlp.up_proj.FLOAT1.000–2
layers.34.mlp.linear_fc2.FLOAT1.000–2
layers.34.adaLN_mlp_ln.FLOAT1.000–2
layers.34.adaLN_sa_ln.FLOAT1.000–2
layers.35.FLOAT1.000–2
layers.35.self_attention.FLOAT1.000–2
layers.35.self_attention.to_q.FLOAT1.000–2
layers.35.self_attention.to_k.FLOAT1.000–2
layers.35.self_attention.to_v.FLOAT1.000–2
layers.35.self_attention.to_out.FLOAT1.000–2
layers.35.self_attention.norm_q.FLOAT1.000–2
layers.35.self_attention.norm_k.FLOAT1.000–2
layers.35.mlp.FLOAT1.000–2
layers.35.mlp.gate_proj.FLOAT1.000–2
layers.35.mlp.up_proj.FLOAT1.000–2
layers.35.mlp.linear_fc2.FLOAT1.000–2
layers.35.adaLN_mlp_ln.FLOAT1.000–2
layers.35.adaLN_sa_ln.FLOAT1.000–2
adaLN_modulation.FLOAT1.000–2
final_norm.FLOAT1.000–2
final_linear.FLOAT1.000–2

Outputs (1)

NameTypeDescription
MODELMODEL