Nucleus-Image
Nucleus-Image is a text-to-image generation model based on Sparse Mixture of Experts (Sparse MoE) released by Nucleus AI on April 14, 2026. It has a total size of 17B parameters but is designed to activate only about 2B during a single inference, taking a different approach from dense models that calculate all parameters each time. It consists of 64 routed experts and shared experts, and uses Expert-Choice Routing to select the tokens that each expert will process.
Nucleus-Image is a text-to-image generation model based on Sparse Mixture of Experts (Sparse MoE) released by Nucleus AI on April 14, 2026. While the total size is 17B parameters, it is designed to activate only about 2B during a single inference, taking a different approach from dense models that calculate all parameters each time. It consists of 64 routed experts and shared experts, and applies Expert-Choice Routing to select the tokens that each expert will process. Instead of deploying the entire large expert organization for each request, it is closer to a system that calls only a few experts related to the task. It supports various aspect ratios, is connected to the Hugging Face Diffusers ecosystem, and is said to provide text KV caching to reduce repetitive text condition calculations.
Existing large-scale image generation models tend to increase inference computation and memory load as the number of parameters increases to improve performance. Furthermore, even if it is an open-source model, if only the final weights are provided and the training code or data composition method is missing, it is difficult for researchers to reproduce the results or verify the model design. The key difference of Nucleus-Image is that it maintains the total capacity of 17B while using a sparse activation structure that calculates only about 2B selected paths for each input, and it also releases the model weights, training code, and data recipe. Like GPT-based models that select appropriate processing paths for text tokens, this model distributes tokens in the image generation process to different experts. Therefore, it is meaningful not only as a simple image generator but also as an experimental subject for studying MoE routing, expert specialization, and reproducibility of training.
Life science researchers can use it to create conceptual images for conference presentations based on paper abstracts or experimental plans. For example, complex scenes such as "interaction between T cells and cancer cells in the tumor microenvironment" can be generated in multiple aspect ratios, and then used in slides or research proposals, clearly stating that it is a synthetic image for explanation rather than actual microscopic data. By combining the Diffusers pipeline and text KV caching, a repetitive generation workflow can be created that changes only the cell type, background, or visual composition in the same long prompt. However, the generated results are not biological facts or experimental evidence, so the structural accuracy should be reviewed separately by experts, and they should not be used for diagnostic, quantitative analysis, or replacement of original data.
Another direction of use is domain adaptation research based on the released weights and training resources. The research team can prepare a separate dataset consisting of life science diagrams or open-source scientific images, and analyze whether the selection patterns of the 64 routed experts differ depending on the tissue type, staining method, or diagram style. By comparing the same prompt with a dense model and recording the generation quality, number of active parameters, latency, and memory usage, the effect of Sparse MoE on scientific visualization can be quantitatively evaluated. The actual fine-tuning procedure, supported parameters, and recommended hardware specifications should be confirmed again in the official repository documentation.
💻 System Requirements
{ram: "To be confirmed", vram: "To be confirmed", storage: "To be confirmed"}
Actual model checkpoint file size needs to be verified.
⚡ Installation
4-1. Quick Start
The installation instructions should be verified from the official GitHub README or the Hugging Face model card. Since the input data does not include verified package names and installation commands, arbitrary pip commands must not be provided.
4-2. Detailed installation
While it has been confirmed that Diffusers-based usage is supported, the code for required dependencies, checkpoint loading, precision settings, and offloading options must be written after rechecking the official documentation.
FAQ
What is Nucleus-Image?
Nucleus-Image is a text-to-image generation model based on Sparse Mixture of Experts (Sparse MoE) released by Nucleus AI on April 14, 2026. While the total size is 17B parameters, it is designed to activate only about 2B during a single inference, taking a different approach from dense models that calculate all parameters each time. It consists of 64 routed experts and shared experts, and applies Expert-Choice Routing to select the tokens that each expert will process. Instead of deploying the entire large expert organization for each request, it is closer to a system that calls only a few experts related to the task. It supports various aspect ratios, is connected to the Hugging Face Diffusers ecosystem, and is said to provide text KV caching to reduce repetitive text condition calculations. Existing large-scale image generation models tend to increase inference computation and memory load as the number of parameters increases to improve performance. Furthermore, even if it is an open-source model, if only the final weights are provided and the training code or data composition method is missing, it is difficult for researchers to reproduce the results or verify the model design. The key difference of Nucleus-Image is that it maintains the total capacity of 17B while using a sparse activation structure that calculates only about 2B selected paths for each input, and it also releases the model weights, training code, and data recipe. Like GPT-based models that select appropriate processing paths for text tokens, this model distributes tokens in the image generation process to different experts. Therefore, it is meaningful not only as a simple image generator but also as an experimental subject for studying MoE routing, expert specialization, and reproducibility of training. Life science researchers can use it to create conceptual images for conference presentations based on paper abstracts or experimental plans. For example, complex scenes such as "interaction between T cells and cancer cells in the tumor microenvironment" can be generated in multiple aspect ratios, and then used in slides or research proposals, clearly stating that it is a synthetic image for explanation rather than actual microscopic data. By combining the Diffusers pipeline and text KV caching, a repetitive generation workflow can be created that changes only the cell type, background, or visual composition in the same long prompt. However, the generated results are not biological facts or experimental evidence, so the structural accuracy should be reviewed separately by experts, and they should not be used for diagnostic, quantitative analysis, or replacement of original data. Another direction of use is domain adaptation research based on the released weights and training resources. The research team can prepare a separate dataset consisting of life science diagrams or open-source scientific images, and analyze whether the selection patterns of the 64 routed experts differ depending on the tissue type, staining method, or diagram style. By comparing the same prompt with a dense model and recording the generation quality, number of active parameters, latency, and memory usage, the effect of Sparse MoE on scientific visualization can be quantitatively evaluated. The actual fine-tuning procedure, supported parameters, and recommended hardware specifications should be confirmed again in the official repository documentation.
When should I use Nucleus-Image?
Nucleus-Image is a text-to-image generation model based on Sparse Mixture of Experts (Sparse MoE) released by Nucleus AI on April 14, 2026. It has a total size of 17B parameters but is designed to activate only about 2B during a single inference, taking a different approach from dense models that calculate all parameters each time. It consists of 64 routed experts and shared experts, and uses Expert-Choice Routing to select the tokens that each expert will process.
📝 Update Notes
No update notes yet.
🧪 Related Code of Life
No related Code of Life posts yet.