Plans & Pricing

Choose How You Build.

Pay as you go for flexible model usage, or subscribe for predictable monthly access with generous limits, priority edge routing, and undrfted. video generation.

Pay-as-you-go

Usage-Based Access

Route requests through Voxvey with usage-based billing and no monthly subscription. This is the fastest way to test workloads, launch pilots, and scale variable usage.

  • Use supported models through one OpenAI-compatible gateway
  • Good fit for experimentation and uneven traffic
  • Usage pricing and setup live on the developer site
Developer Site

Compare Subscription Tiers.

Compare monthly access, rate limits, video generation allowances, and support options across every Voxvey subscription tier.

Feature Free $0/mo Popular Pro $30/mo Enterprise $300/mo
Model Access Select open-weight All 350+ models All 350+ models
Usage-Based Billing
Chat Playground Unlimited Unlimited
API Rate Limits Basic Generous Highest available
undrfted. Site Generations 3 per week 30 per month
Edge Routing Priority Standard Priority Dedicated capacity
Support Community Email (24hr) Dedicated account manager
Custom SLA
Consolidated Invoicing

Frequently Asked Questions.

What models are available on the Free tier?

The Free tier includes access to select open-weight models like Llama 3.3, DeepSeek V3, and Qwen 2.5. Premium closed-source models like GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro require a Pro or Enterprise subscription.

Do unused video generations roll over?

Video generation allowances reset each billing cycle. Pro plans receive 3 generations per week (resetting weekly), and Enterprise plans receive 30 generations per month (resetting monthly). Unused generations do not carry over.

Can I cancel my subscription at any time?

Yes. There are no long-term contracts. You can upgrade, downgrade, or cancel your subscription at any time from the developer console.

What is priority edge routing?

Voxvey's API runs on a globally distributed edge network with 300+ locations. Pro and Enterprise subscribers receive priority placement in the dispatch queue, resulting in lower latency and faster time-to-first-token on every request.