Read the Code in This Order
The files are split by runtime responsibility. For a new model, read the request path first:configs/pipeline_configs/{model}.pydefines model-specific denoising and decoding behavior.runtime/pipelines/{model}.pywires modules into stages.runtime/pipelines_core/stages/runs the shared stage logic.runtime/models/contains native model components only when the architecture cannot be reused.
runtime/models/ owns modeling code: checkpoint-defined neural modules,
architecture wrappers, and weight-loading or forward-path details that are
intrinsic to one model family. Reusable serving infrastructure belongs in
SGLang-Diffusion runtime folders such as runtime/cache/,
runtime/distributed/, runtime/utils/, or shared pipeline stages. This
includes cache managers, graph runners, process-group transport, request
utilities, and common action-policy helpers. Model packages may call these
helpers. Keep ownership in shared runtime folders unless the code is truly
architecture-specific.
Place helpers with their owners
Use the narrowest existing owner before adding a utility module:
Do not create a top-level
utils/ package or grow a catch-all utils.py or
common.py. Split large mixed-responsibility files along ownership boundaries,
not arbitrary line counts. A helper folder is warranted only when several
cohesive modules need it, not for a single function or hypothetical reuse.
Model code must not import pipeline stages. Put contracts shared by models and
stages in a lower-level domain module; for example, realtime cache keys belong
under runtime/realtime/. Keep GPU initialization, monkey patches, and model
loading out of generic utility imports.
When moving internal helpers, update all callers, tests, and cookbook examples
together. Preserve documented registration and serving entry points; do not
add re-export chains just to retain obsolete internal utility paths.
Out-of-Tree Models and Pipelines
An installed package can register native component models and a pipeline without modifying SGLang-Diffusion. Register them in the package’s__init__.py:
- The string form of
register_modelkeeps component imports lazy. hf_model_pathsalso supports checkpoints withoutmodel_index.json. Other Diffusers checkpoints can select the pipeline through_class_name.- For a standalone safetensors file, pass
--pipeline CustomPipeline. - Set the environment variable before startup. Each process imports the package
once. Use
overwrite=Trueonly to intentionally replace a built-in pipeline.
Start With the Smallest Change
Before adding files, decide which path fits the model.
Do not add a folder just to mirror the Diffusers repository layout. Add a new
file only when an existing pipeline, stage, module, config, or sampler cannot
express the behavior clearly.
Minimal File Map
The source tree is split by runtime responsibility. That split is useful for optimization. Keep new model PRs focused on the files required by model behavior.
For a new native architecture, the common minimum is:
configs/sample/{model}.pyconfigs/pipeline_configs/{model}.pyruntime/pipelines/{model}.pyruntime/models/dits/{model}.py
runtime/models/. Prefix caches,
request-local contexts, denoising graph runners, OpenPI-compatible transport,
and prefix/action process-group utilities should be shared SGLang-Diffusion
runtime infrastructure when they are useful beyond the first model.
Read the Reference First
Use the model’s Diffusers pipeline, official implementation, ormodel_index.json as the source of truth. Write down:
- Which modules must be loaded: tokenizer, text encoder, image encoder, transformer, scheduler, VAE, processor, and any extra adapters.
- The prompt and image encoding flow.
- Latent shape, packing, scale, shift, dtype, and device rules.
- Timestep and sigma schedule.
- The exact
forward()kwargs expected by the denoising network. - VAE decode rules and output post-processing.
Conditioning reuse
Use the nativeTextEncoder / ImageEncoder base classes (or EncoderTensorParallelMixin) and call encoders through encoder(...). This preserves the shared conditioning cache, module hooks, and encoder TP/folding group. A replicated custom encoder can use ConditioningEncoderMixin; do not apply it to a module with hidden mutable state or random sampling in forward().
The standard TextEncodingStage caches postprocessed embeddings, conditioning masks, sequence lengths, and pooled outputs instead of unused intermediate hidden states. Keep its postprocessing deterministic. When adding a stage-level cached_encoder_call, place deterministic preprocessing and component preparation inside the compute callback and give the stage a namespace. The namespace avoids a duplicate full encoder entry while preserving nested vision-method caches. Declare conditional component uses with start_at_stage_entry=False, so the executor cannot load weights before checking the cache. Keep calling encoder(...) inside the callback to preserve hooks and parallel-group handling. After retrieving the conditioning, call finish_unused_declared_component() to release weights retained by warmup or the previous request when no encoder use began; an active use remains owned by the residency manager. Batch-DP retains per-copy caches. A gathered group result may skip encoding and the output gather only after hit consensus across the whole replica; a miss keeps every rank in the normal gather path.
For a custom VAE, apply cached_vae_encode only to the deterministic encode method returning a tensor or posterior. Keep posterior sampling, normalization, initial noise, scheduler state, and request-specific KV dictionaries outside that boundary. Unsupported outputs bypass caching; explicitly register a tensor-only posterior container with register_conditioning_container when needed. Never register an autoregressive cache or a session object as a posterior container.
For integrated models without a separate encoder, use cached_conditioning on a pure conditioning method, as in Cosmos3’s understanding pathway or SenseNova-U1’s reference-image features. Every rank in the current encoder TP group must enter that method together. Do not decorate the noisy-image or denoising path.
Prefer these function boundaries to whole-stage deduplication. A stage may also write per-request metadata, advance RNG state, release materials, or execute collectives; these actions must still run for each request. The executor gives each grouped or batched stage a temporary device-result scope, using the same keys and invalidation as the cross-request cache. Use share_in_group=True only for consumed tensors that remain read-only across requests; containers are copied. Other outputs get private tensor snapshots with a bounded temporary budget. The scope ends at stage exit, including exceptions, and remains available with --disable-conditioning-cache.
Wrap negative encoding in prefer_conditioning_cache() to prioritize reusable conditioning. At a consumed cached_encoder_call boundary with a stage namespace and share_in_group=True, this keeps a private device snapshot across requests, including library encoders and gathered batch-DP results. Hits skip encoder preparation and postprocessing, with the same rank consensus and invalidation as CPU entries. Device and CPU entries share one byte and entry budget; raw intermediate encoder states remain on the CPU. Avoid adding a separate model-owned negative cache.
Do not put Req, scheduler instances, sampled latents, or session state into a cross-request cache. Keep model-specific grouped execution for parallel work distribution and material ownership. Preserve batch shapes and padding when matching batched encoder calls; per-row regrouping needs separate numerical validation.
The common TextEncodingStage retains stage deduplication to avoid repeated tokenization and stage bookkeeping. If you override its forward, declare a complete stage-output contract or set deduplicated_output_fields = () and use conditioning-result reuse. Do not inherit the common stage’s output list when your stage also writes model-specific metadata. Keep request-local containers independent even when their tensors are shared.
Test cache-on versus cache-off with identical seeds, changed prompts/masks/reference pixels (including alpha), downstream in-place mutations, changed weights, and rank-local eviction under supported parallel groups. Also exercise grouped requests through the executor with a zero or undersized cache budget: check encoder call counts, per-request metadata, output ownership, and group cleanup on failure. Include the new model in the conditioning cache coverage.
Preserve existing grouped-request encoding reuse when cross-request caching is unavailable. Use cross_request=False for boundaries that support only group-local reuse, such as library fallback encoders and gathered batch-DP outputs. FSDP scopes allow consumed-output group reuse while bypassing nested encoder caches. A boundary containing collectives needs hit consensus across every participating rank, including FSDP shards and encoder DP copies; an encoder TP group alone may be insufficient. Test these bypass paths and single-rank invalidation before removing stage deduplication.
Choose a Pipeline Shape
SGLang-Diffusion usesComposedPipelineBase to wire stages together. Most
native pipelines should choose the least invasive stage shape that preserves the
runtime semantics.
Prefer this order:
- Use native stages directly. This keeps the model on shared code paths for offload, component readiness, profiling, disaggregation, batching, and future stage-level optimizations.
- Subclass the narrowest native stage. If only prompt processing differs,
inherit from
TextEncodingStage. If only latent setup, timestep setup, denoising, or decode differs, inherit from that specific native stage. Preserve the existing input/output fields whenever possible. - Add a custom single-purpose stage only when no native stage contract fits. Keep the stage owner narrow: one stage should own one coherent transformation, such as a custom condition assembly step or a model-specific policy/action bridge.
- Use an aggregated
BeforeDenoisingStageonly as a last resort. This is the least preferred shape because it hides multiple runtime responsibilities in one stage, increases code size and review cost, and bypasses shared hooks for offload, profiling, disaggregation, batching, and future stage-level optimizations.
Implement the Pieces
1. Sampling Params
Create request parameters only for values users can set at runtime.2. Pipeline Config
PipelineConfig is where shared denoising and decoding stages get model-specific
callbacks.
forward() signature exactly.
3. Pipeline Wiring
Use the standard helper when the model fits it.TextEncodingStage directly.
BeforeDenoisingStage only when the reference pipeline couples
several preparation steps so tightly that splitting them would require fragile
duplicate state or extra synchronization. Do not start with this shape.
Reuse component loaders
Most components should keep the default loader for their role (transformer, text encoder, VAE, scheduler, or tokenizer). For an auxiliary module that acceptsmodel_cls(**config) and loads an unchanged state dict, select
PlainStateDictComponentLoader instead of adding a model-specific loader class:
_extra_config_module_map. It is local to this pipeline; omitted components keep
their existing dispatch. An explicitly selected loader fails on errors instead
of falling back to a different implementation.
Register each module through EntryClass in its model file, or
ModelRegistry.register_model for an out-of-tree package. The shared loader:
- Resolves the registered class from
config.json’s_class_name, falling back to the architecture inmodel_index.json, and passes non-metadata config fields to its constructor. - Loads a single safetensors file or an indexed sharded checkpoint using the shared checkpoint selector, with strict key and shape checks.
- Honors exact overrides such as
--component-weights-paths.queryformer PATHand--component-precisions.queryformer bf16. Precision defaults to the pipeline’sdit_precision. - Leaves eval mode and final CPU/GPU placement to the shared loading lifecycle. It does not implement TP/FSDP weight sharding, quantization, or direct GPU loading. Explicit quantization and direct-GPU-loading overrides are rejected.
component_loaders; registering a model
alone does not opt it into the plain state-dict protocol.
4. Last-Resort Before-Denoising Stage
ABeforeDenoisingStage is not a catch-all replacement for the native stages.
Use it when the model has custom latent packing, conditioning assembly, timestep
preparation, or request-local state that does not fit LatentPreparationStage or
TimestepPreparationStage, and only after checking whether the work can be a
native-stage subclass or a custom single-purpose stage. If the difference is
prompt handling, subclass TextEncodingStage instead.
A proper BeforeDenoisingStage should populate the batch fields consumed by
DenoisingStage.
DenoisingStage:
5. Distributed and memory integration
Single-GPU parity is only the first milestone. Complete native support also requires:- Encoder and DiT TP/SP: use native parallel projections and sharded weight
loading for TP, and
USPAttentionfor SP. Handle masks, RoPE, padding, and output gathering without falling back to a replicated full model. TP and SP must work together. - VAE parallel decode: subclass
ParallelTiledVAE, or reuse an existing native base with the same contract. Support tiled andspatial_sharddecode throughDecodingStageand the shared decode group. Reuseruntime/layers/parallel_conv.pyandruntime/models/vaes/parallel/diffusers_spatial.pywhere applicable. - Layerwise offload: every loaded neural module must inherit
LayerwiseOffloadableModuleMixinand list all repeated block paths inlayer_names. Setlayerwise_offload_dit_group_enabled = Falsefor non-DiT modules. Component CPU offload is not a substitute.
wanvideo.py and qwen_image.py for DiT TP/SP, gemma_3.py for encoder TP
and offload, and autoencoder_kl_qwenimage.py or ltx_2_vae.py for VAE decode.
The Diffusers backend is compatibility-first and does not need to meet this
native integration contract.
6. Registry
Define aregister() function in configs/pipeline_configs/{model}.py. The
runtime auto-discovers it on startup and calls it to register the sampling
params and pipeline config.
model_detectors matches a model path or model_index.json _class_name when
the Hugging Face path varies; see wan.py or qwen_image21.py for real
examples. The pipeline file is discovered through its EntryClass; do not add
a second pipeline registry unless the existing registry requires it.
Verify the Port
Use one deterministic prompt and seed while comparing with the reference implementation.- Run a single-GPU smoke test and check that the output contains coherent content.
- Compare latent scale and shift, timestep order, sigma values, and conditioning kwargs against Diffusers or the official implementation.
- Compare VAE decode separately, including tiled and multi-GPU
spatial_shard. - Run encoder and DiT TP, SP, combined TP x SP, and
--layerwise-offload-components all; compare with the single-GPU resident baseline. - If the model supports LoRA, CFG parallelism, or disaggregation, test each feature explicitly.
- Add or update the cookbook, examples, and Supported Models catalog when users need a new launch command.
- Wrong latent scale or shift.
- Reversed or dtype-mismatched timesteps.
- Missing negative embeddings when CFG is enabled.
- Conditioning kwarg names mismatched with the DiT
forward(). - Rotary embedding shape or style mismatch.
- Decoding packed latents without restoring
raw_latent_shape.
PR Checklist
- Reused an existing family, stage, module, scheduler, or VAE wherever possible.
- Kept the new-model touch surface small and justified any extra files.
- Added
SamplingParams,PipelineConfig, pipeline wiring, DiT module, and registry entry when native support is needed. - Confirmed
pipeline_namematches the Diffusersmodel_index.json_class_namewhen applicable. - Confirmed
_required_config_modulesmatches the model repo. - Verified image or video quality against a reference output.
- Completed the distributed and memory integration checks above.
- Tested CFG parallelism and distributed serving paths when they apply.
Declare multiple task types
KeepPipelineConfig.task_type as the default task. To let one configured
pipeline serve several tasks, declare its capabilities as a class variable:
(task_type,). Override get_supported_task_types() if capabilities depend
on the loaded checkpoint or configuration, preserving these invariants.
Flat CLI/Python task_type arguments select the request. To change the
configured default, use pipeline_config={"task_type": ModelTaskType.I2V}
(or its serialized configuration), choosing a supported task.
Capabilities describe implemented execution paths; declaring a task does not
load additional components or implement conditioning and decoding for it.
Requests carry SamplingParams.task_type, accepting an enum or its name
(case-insensitive). Admission resolves it to an enum and derives the output
data_type. Explicit tasks must belong to the supported set and match the
HTTP endpoint’s output type. Without an explicit task, admission filters by
output type and conditioning inputs, prefers a compatible default, then
selects a unique compatible alternative. Python/CLI requests also prefer the
default’s output type among alternatives. Ambiguous requests must specify
task_type. No request changes the shared configuration.
Use the admitted batch.task_type in model stages when behavior depends on
the request. Shared stages use get_request_task_type(batch, pipeline_config)
to preserve compatibility with direct callers that have no resolved task.
Audit model-specific frame adjustment, reference loading, latent construction,
and decoding for assumptions about the server default. Declare the union of
components used by all supported tasks in stage residency metadata. Synthetic
warmup uses the default task; provide the pipeline’s warmup hook when that task
needs model-specific reference inputs. Dynamic batching compares the admitted
task along with the other sampling fields.
I2I and I2V require image input; T2I and T2V reject it. Existing TI2I
and TI2V retain optional image conditioning. V2V and F2V require video
input; F2V denotes video-prefix conditioning. Models own the details of
prefix length and temporal alignment. Pipelines with explicit capabilities
reject video input for tasks that do not accept it; legacy singleton pipelines
retain their model-owned video checks.
See request task selection
for the HTTP interface. Test task selection, invalid inputs, discovery,
serialization, batching, and alternating tasks on one running server.