(LoRA) Fine-Tuning FLUX.1-dev on Consumer Hardware
Related stories
Using LoRA for Efficient Stable Diffusion Fine-Tuning
ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs
arXiv:2609.06072v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. This expert-w...
Beyond LoRA: Can you beat the most popular fine-tuning technique?
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
arXiv:2603. 16428v2 Announce Type: replace-cross Abstract: Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabilities of most GPUs.
StateTune: Transforming LLM-Assisted EDA Flow Tuning into a Stateful, Closed-Loop Process
arXiv:2608.23601v1 Announce Type: cross Abstract: EDA flow parameter tuning is critical for quality-of-results~(QoR), yet the parameter space is large, tightly coupled, and full evaluations are prohi...
DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs
arXiv:2607. 28848v1 Announce Type: cross Abstract: LLM serving systems are provisioned for peak load to meet strict latency targets, leaving substantial GPU compute idle whenever traffic falls below peak.
JET: Justification Evaluation in Transformer
JET (Justification Evaluation in Transformer) leverages pretrained language and vision‑language models to choose among a limited set of answers without extra training. It directly evaluates candidate likelihoods, reuses computation across candidates, and runs experiments on desktop CPUs and consumer GPUs to measure decision accuracy and execution cost. Results show high accuracy on the MMLU test set, significant speedups from prefix reuse and cache management, and a 30.8% reduction in process time through input preparation optimizations, all while maintaining unchanged outputs.
FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving
FluxMoE introduces an expert paging system that decouples Mixture-of-Experts (MoE) model experts from permanent GPU residency, allowing dynamic adaptation to available memory. By combining PagedTensor, a bandwidth‑balanced memory hierarchy, and a budget‑aware residency planner, FluxMoE streams expert weights on demand while keeping computations on GPUs. Experiments on GLM‑4.5 and Mixtral‑8×7B‑Instruct show significant throughput gains and reduced time‑per‑output‑token compared to existing inference engines, without compromising model quality.
Understanding LoRA Rank Trade-offs in Diffusion Model Fine-Tuning
The study investigates how the rank of Low‑Rank Adaptation (LoRA) affects diffusion model fine‑tuning on CIFAR‑10 using a DDPM U‑Net. Experiments with ranks 2, 4, 8, 16, and 32 show that moderate ranks—particularly rank 4—yield the best FID scores while keeping trainable parameters, runtime, and GPU memory low. Higher ranks offer only marginal improvements at a higher computational cost, suggesting that small‑to‑moderate ranks are efficient defaults for fixed training budgets.
Faster assisted generation support for Intel Gaudi
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
arXiv:2609.01343v1 Announce Type: new Abstract: Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating archit...