NOVA ANIME AM ↔ QWEN-IMAGE-2.1
One character,
two models
An archived experiment in translating clean latent estimates between image models.
This learned bridge crosses directly between latents at inference. Its reference targets were prepared by decoding the source and encoding it into the other model. We are preserving these results before investigating target preparation and training that stay entirely in latent space.
These are development results. Test data remains reserved, and there is no matched fixed-cloud result establishing a particle advantage. Output differences do not by themselves rank image quality.
Combined character sampling
One pinned checkpoint pairStart in either model, translate the current clean estimate through the bridge, carry the noise across, and continue denoising. Compare a halfway handoff with a switch and return through both bridges.
Same prompt, seed, shared noise grid and 24-step budget within each row. Full-codec references use the identical switch schedule. All eight learned switching trajectories repeated bitwise exactly. These adult character prompts and seeds are outside all 4,096 planned training prompts and were not used to select checkpoints.
Individual bridge reconstruction
Latest saved validationOne bridge forward pass followed by the recipient decoder, with no further denoising. The kingfisher is held-out validation, used in checkpoint selection; it is not an untouched final-test example.
Learning curves
Additional updates in the resumed, growing-data run. Validation examples stay fixed while training prompts are admitted in blocks of 32. Character images above use their explicitly pinned earlier snapshot.
Validation latent NMSE ↓
LPIPS measures decoded image differences. Latent NMSE measures prediction error after fixed target normalization; it is not a percentage of lost image quality. Switches also change trajectories when using full codecs, so compare learned switching against its matching reference.