Pricing

Free

$0

For getting started with open models

Includes:

  • Run models locally
  • Starter usage credits included
  • Includes access to starter models
  • Add credits to unlock all models
  • No service fees

Pro

$20 / mo. or $200/yr billed annually

For shorter, well-defined day-to-day tasks

Everything in Free, plus:

  • $60 of usage credits per month
  • Access to larger pro models
  • Run multiple models concurrently
  • Fast mode (coming soon)

Team

Early access
$500 / mo.

For teams scaling with open models

Everything in Pro, plus:

  • Unlimited users
  • $1,000 of usage credits per month, shared across the team
  • Centralized billing and administration
  • Priority support
  • Shared projects, skills, and instructions (coming soon)

Enterprise

Custom

For larger teams with volume usage pricing

Everything in Team, plus:

  • Model access controls Limit team access to specific models
  • Set cost budgets for users and API keys
  • Private Slack channel with dedicated support
  • Custom security questionnaires

Model pricing

Prices are per million tokens

Model Input Cached input Output
deepseek-v4-flash $0.44 $0.014 $1.32
deepseek-v4-pro $1.32 $0.044 $3.96
gemma4 $0.14 $0.05 $0.40
glm-5.3 $1.40 $0.26 $4.40
glm-5.3-flash $0.15 $0.03 $0.50
glm-5.2 $1.40 $0.26 $4.40
glm-5.1 $1.00 $0.20 $3.20
gpt-oss:120b $0.15 $0.014 $0.60
gpt-oss:20b $0.07 $0.035 $0.30
kimi-k3 $3.00 $0.30 $15.00
kimi-k2.7-code $0.95 $0.19 $4.00
kimi-k2.6 $0.95 $0.16 $4.00
minimax-m3 $0.60 $0.12 $2.40
minimax-m2.7 $0.30 $0.06 $1.20
mistral-large-3 $0.50 $0.50 $1.50
nemotron-3-nano $0.06 $0.06 $0.24
nemotron-3-super $0.015 $0.015 $0.60
nemotron-3-ultra $0.10 $0.10 $3.00
qwen3.5:397b $0.60 $0.60 $3.60

Frequently asked questions

Team

  • Who is the Team plan for?

    The Team plan is for teams that want shared usage, billing, and administration across the whole team.

  • How does Team billing work?

    Team starts at $500 per month with unlimited users and $1,000 of usage credits included each month, shared across the team.

    Usage beyond the included credits draws from a shared team balance billed as you go. You can add a set amount to the balance and turn off automatic usage billing.

  • Does joining a team change my personal account?

    Ollama creates a separate team account for you when you join. You can switch between your team and personal accounts with the same login.

Models

  • Which models are available?

    See the full list of cloud-enabled models here.

  • Do models support tool calling?

    Yes. Cloud models that are trained to support tools are tested for tool calling and with real agent workflows before they go live. If something isn't working, let us know at support@ollama.com.

  • What quantization or data format do cloud models use?

    Native weights, as released by the model provider. On modern NVIDIA hardware, models may use accelerated data formats supported by Blackwell and Vera Rubin architectures (e.g. NVFP4).

  • How fast is Ollama?

    Speed depends on model size, architecture, and hardware optimization. We target and monitor for low time-to-first-token and high throughput across all cloud models. Priority tiers with faster performance may be available in the future.

Usage

  • How does monthly usage work?

    Pro, Max, and Team plans include a set dollar amount of usage credits each month. Free accounts include a starter amount of usage for a smaller set of starter models. Adding extra credits unlocks all models. Running models on your own hardware is always unlimited.

  • When does my monthly included usage reset?

    On Pro, Max, and Team plans, included usage resets monthly on the same day of the month your subscription started, including on annual plans. On the Free plan, usage resets monthly from the date you signed up.

  • Does unused included usage roll over to the next month?

    No. Instead, your included amount refreshes at each monthly reset.

  • How is usage measured?

    Usage is measured in tokens at each model's rates. The model pricing table lists the input, cached input, and output price per million tokens for every cloud model.

  • How does extra usage work?

    Every plan, including Free, can add extra usage credits. Ollama uses included plan credits first, then draws from the extra usage balance. Team usage draws from one balance shared by the organization.

  • How do I know when I've hit my limit?

    Check your usage here anytime. On paid plans, Ollama sends an email reminder at 90% of your included monthly usage. You can turn this off in settings.

  • How many requests can I run at once?

    Concurrency limits ensure dedicated capacity for workflows that need multiple requests running simultaneously. Free includes 1 concurrent request, Pro 3, and Max and Team 10.

    Requests beyond your plan's concurrency limit are queued and processed as soon as a slot is available. Queued requests are held up to a fixed limit - if the queue is full, the request will be rejected until one of your concurrency slots opens.

  • What happens to my usage if I switch from an existing Pro or Max plan to the new one?

    Your usage is reset: the new plan's full monthly amount is available as soon as you switch, and the session and weekly limits of the old plans no longer apply. Your reset date stays on your subscription's original start date. Switch any time from your billing settings.

Accounts

  • Can I have multiple Ollama accounts?

    No. Ollama is one account per person.

Privacy

  • Where are models hosted?

    Ollama hosts models and compute resources primarily in the United States. To serve global demand, we may route to Europe and Singapore for additional capacity.

  • Is my prompt or response data trained on?

    Prompt or response data is never logged or trained on.

  • Who does Ollama partner with to host models?

    Ollama collaborates with NVIDIA Cloud Providers (NCPs) to host open models.

    When Ollama partners with providers, we require no logging, no training, and zero data retention policies in place.