Fireworks AI

Fireworks AI - Coding & Development AI Tool

Inference and fine-tuning at published rates.

Fireworks serves open models through a fast inference API and handles LoRA and full-parameter fine-tuning\, with training priced per million tokens and dedicated GPUs available by the hour.

Publishing training rates by method and model size is genuinely useful — most vendors make you ask. LoRA SFT on a model up to 16B is $0.50 per million training tokens\, which makes experimentation affordable rather than a budget conversation.

Quick Information

Platform

Web

Pricing

Usage-based + $1 free credits

API

Available

Category

Coding & Development

Pros and Cons

Pros

  • Published fine-tuning rates by method
  • Fast serverless inference
  • Dedicated GPU deployments available
  • High rate limits
  • Affordable LoRA experimentation

Cons

  • Only $1 in free credits
  • GPU hourly rates rising
  • Open models only
  • Pricing structure is complex
$0.00

Usage-based + $1 free credits

Serve Open Models Fast

Fine-Tune With LoRA Cheaply

Rent Dedicated GPUs

Generate Embeddings at Scale

FAQs

How much does fine-tuning cost?

For models up to 16B: LoRA SFT $0.50\, LoRA DPO $1.00\, full-parameter SFT $1.00 and DPO $2.00 per million training tokens. Rates scale up with model size.

Is there a free tier?

You "get started with $1 in free credits" on serverless inference with postpaid billing.

What do dedicated GPUs cost?

On-demand H100 and H200 deployments are listed at $7.00/hour\, B200 at $10.00 and B300/GB300 at $12.00\, with announced increases from September.

What about embeddings?

Models up to 150M parameters are $0.008 per 1M input tokens\, 150M–350M are $0.016.

Write Your Own Review

Write Your Own Review