The Next Layer of Intelligence
Spectra-FFT trains large models in the frequency domain, cutting optimizer memory by over 50% on the GPUs you already own.
Get Started View Architecture
Get Started
Spectra-FFT trains large models in the frequency domain, cutting optimizer memory by over 50% on the GPUs you already own.
Get Started View ArchitectureThe problem
AdamW keeps two 32-bit states for every parameter: 8 bytes each, on top of the weights. That state is what pushes full-parameter fine-tuning into out-of-memory crashes on constrained hardware. Spectra-FFT is a new optimizer that stores far less of it.
Gradients carry signal and noise. Spectra works in the frequency domain, where the two separate cleanly.
Optimizer state is stored for the part of the signal that drives learning, not for every parameter.
Full-parameter training with no low-rank adapters and a much smaller memory footprint.
Architecture
Spectra-FFT sits exactly where your optimizer does today. Nothing upstream or downstream changes: same model, same data, same loss, same scheduler.
One line replaces AdamW(...) with Spectra(...). Works with your existing training scripts, mixed precision and checkpointing.
Instead of two full-size states per parameter, Spectra maintains a compact representation of the training signal, which is where the 50%+ optimizer-memory reduction comes from.
Runs entirely on-device on CUDA using standard NVIDIA libraries. No CPU offloading, no custom hardware, no extra data movement.
Every weight is trained. This is not LoRA or an adapter, so there is no low-rank ceiling on what the model can learn.
The concept, the measured results, the paper and the training logs. Everything needed to evaluate whether Spectra works.
The implementation, parameters and tuning that make it work. Available to evaluation partners under NDA.
VRAM engine
Estimate assumes bf16/fp16 weights (2 bytes/param) and fp32 AdamW momentum + variance (8 bytes/param). Spectra figure applies the ~55% optimizer-state reduction reported in our paper. Activations, gradients and KV cache are excluded and depend on batch size and sequence length.
Evidence
Full-parameter fine-tuning inside a strict 5GB VRAM partition of an NVIDIA A100, logged end to end.
Evaluation access
We share a private Google Colab notebook that trains the same model twice on a standard GPU: once with AdamW, which runs out of memory, and once with Spectra-FFT, which completes. Access is granted to engineers and teams evaluating Spectra.
# What the notebook demonstrates opt = AdamW(model.parameters()) # -> CUDA out of memory opt = Spectra(model.parameters()) # -> trains within budget
Requests are reviewed manually, usually within 48 hours. Evaluation access is provided under a short NDA.
FAQ
Yes. It is a drop-in optimizer for full-parameter training, aimed at memory-constrained GPUs.
Not currently. The paper describes the method and the demo lets you verify the behavior. Evaluation and partnership access is available on request.
Our logged runs track baseline convergence on the evaluated setup. See the paper and W&B logs for the exact configuration and limits.
Teams fine-tuning models on-premise or on consumer and mid-range GPUs, where optimizer state is the bottleneck.