Direct Answer: To run FLUX.1 without CUDA out-of-memory errors on 12GB–16GB consumer GPUs, load FLUX.1 dev in FP8 (or GGUF Q4_K_M / Q8_0) using ComfyUI’s native UNETLoader and DualCLIPLoader, offload CLIP and T5 to CPU RAM, and clamp inference to 20–25 steps with Euler sampler. Only introduce LoRA nodes after validating this baseline.
Start here
Intended reader: Creators running local RTX 3060/4070/4080 GPUs seeking photorealistic outputs from FLUX.1. Practical outcome: A crash-free node setup that loads weights efficiently and accepts LoRAs without breaking latent dimensions. Your base image looks promising, but the moment you add a LoRA, several other things change too. Was it the model, the prompt or the new component? Build a baseline you can return to, then make one deliberate change.
Choose the model deliberately
Black Forest Labs identifies FLUX.1 dev as a 12-billion-parameter model. Its model card links a specific non-commercial model license and separately describes permitted output uses. FLUX.1 schnell lists Apache 2.0. Treat model licensing and output usage as separate checks; read the terms for the exact files you download.
Begin with the official inference example
ComfyUI's FLUX.1 tutorial supplies a workflow and explains the model components. Follow its file placement instructions, load the graph and confirm that selected files match the example. An inference workflow generates an image from existing weights; it does not train a new LoRA.
Save a reproducible baseline
Use one prompt and fixed settings until a complete image succeeds. Store the graph alongside a short environment note. This makes it easier to separate a missing file from a creative change or incompatible custom node.
- Record ComfyUI version and selected model filenames.
- Save prompt, seed, dimensions and sampling settings.
- Record peak memory and elapsed time on your own hardware.
- Keep the first working output as a reference.
- Download clip_l.safetensors and t5xxl_fp8_e4m3fn.safetensors into ComfyUI/models/clip/.
- Download flux1-dev.safetensors into ComfyUI/models/unet/ or models/checkpoints/.
- Download ae.safetensors into ComfyUI/models/vae/.
- Connect DualCLIPLoader (clip_l + t5xxl) to CLIPTextEncode with clip_type set to "flux".
- Set KSampler: steps=20, cfg=1.0, sampler_name=euler, scheduler=simple.
- Run a single test generation at 1024x1024 to verify VRAM allocation before loading LoRA weights.
Add a LoRA as one controlled change
Add one compatible LoRA, compare it with the baseline and record its strength. Check the author's base-model requirements. Avoid stacking unfamiliar adapters while debugging. For training, treat dataset preparation, trainer compatibility and evaluation images as a separate project. A working generation graph does not validate a training configuration.
Memory savings need a measured trade-off
The FLUX.1 dev example demonstrates CPU offloading to reduce GPU memory use. That does not establish a universal minimum VRAM figure. Dimensions, extra nodes, precision and other resident workloads change what fits. Compare a small fixed job before and after a memory-saving change; preserve its output and timing.
| Model Variant | File Size / Precision | Min VRAM Required | System RAM Requirement | Licensing & Best Fit |
|---|---|---|---|---|
| FLUX.1 [schnell] | ~23.8 GB (FP16) or ~12 GB (FP8) | 12 GB VRAM (FP8) | 32 GB DDR4/DDR5 | Apache 2.0 (Commercial allowed) · Fast 4-step preview |
| FLUX.1 [dev] | ~23.8 GB (FP16) or ~11.9 GB (FP8) | 16 GB VRAM (FP8/NF4) | 32–64 GB DDR5 | Non-commercial research · High-fidelity 20-30 steps |
| FLUX.1 GGUF (Q4_K_M) | ~6.5 GB (4-bit quantized) | 8–10 GB VRAM | 24 GB System RAM | Community quantized · Budget GPUs (RTX 3060/4060) |
| FLUX.1 [pro] | Cloud API Only | N/A (Managed API) | N/A | Commercial paid API · When local hardware is unavailable |
A useful stopping point
Your first milestone is a saved graph that another person can reopen with the same files and understand. Add a short note describing what each extra adapter contributes. Only then scale dimensions, batch size or training complexity. This guide is documentation-based; it does not claim a tested speedup on a particular GPU.
Keep an experiment card beside the workflow
Our suggested experiment card has six fields: model files, workflow version, input or prompt, changed setting, expected effect and observed result. Fill it in before introducing a LoRA. This makes it harder to mistake a different prompt, seed or resolution for the effect of the LoRA itself. A fixed seed is helpful within a controlled setup, but does not guarantee identical outputs after hardware or software changes. Preserve the complete baseline rather than only a screenshot of the nodes.
Separate inference from a training decision
Using an existing LoRA and training your own are different projects. Before training, define the visual behavior you cannot obtain from your current workflow and collect example material you have permission to use. Reserve some examples for evaluation rather than treating every pleasing result as proof of generalization. This is a planning recommendation; the exact data preparation and training settings belong to the selected training tool and model license. Do not start with a promised universal rank or step count.
Choose the next change by the problem you can see
If the image is compositionally wrong, investigate the input and conditioning before optimizing memory. If the workflow fails to load, verify required files and component compatibility before rewriting the prompt. If the result changes after an update, compare the saved baseline on the new environment before adding more nodes. Solve one category of failure at a time. The useful milestone is not the longest graph: it is an output you can explain, reproduce closely enough for the project and hand off with clear notes.
Try this next
If the graph still feels unfamiliar, start with the first-image walkthrough. Read ComfyUI beginner guide: your first image and common fixes.
Sources & further reading
Primary sources checked Sep 12, 2026. Vendor statements are attributed; editorial advice is our own.
- 1FLUX.1 dev model card ↗Black Forest Labs
- 2FLUX.1 schnell model card ↗Black Forest Labs
- 3FLUX.1 text-to-image workflow ↗Comfy Org
Help us keep this useful. Send a correction or a primary source →



