Inference and fine-tuning at published rates.
Fireworks serves open models through a fast inference API and handles LoRA and full-parameter fine-tuning\, with training priced per million tokens and dedicated GPUs available by the hour.
Publishing training rates by method and model size is genuinely useful — most vendors make you ask. LoRA SFT on a model up to 16B is $0.50 per million training tokens\, which makes experimentation affordable rather than a budget conversation.
Quick Information
Platform
Web
Pricing
Usage-based + $1 free credits
API
Available
Category
Coding & Development
Pros and Cons
Pros
- Published fine-tuning rates by method
- Fast serverless inference
- Dedicated GPU deployments available
- High rate limits
- Affordable LoRA experimentation
Cons
- Only $1 in free credits
- GPU hourly rates rising
- Open models only
- Pricing structure is complex

