Skip to main content

Overview

These are the models supported by Cloudidr LLM Ops to track API costs. The model pricing is from the providers which we use to calculate your spend.
Up to date supported models, providers and pricing can be found in the LLM Ops left side panel Starting GuideModel Pricing tab
Last Updated: January 10, 2026Model pricing is subject to change by the providers. We update our pricing regularly to ensure accurate cost tracking.
Model Not Listed?If your model is not in this list, please contact us at support@cloudidr.com and we’ll add support for it.

Pricing Tables

Anthropic Claude Models

All pricing is per 1 million tokens.
Model Recommendations:
  • Opus - Most capable, best for complex reasoning
  • Sonnet - Balanced performance and cost
  • Haiku - Fastest and most affordable

Integration Guide

See the Anthropic Integration page to start tracking costs.

Cost Comparison

Best for high-volume, simple tasks:
  • GPT-5 Nano: 0.05input/0.05 input / 0.40 output
  • Gemini 2.0 Flash Lite: 0.075input/0.075 input / 0.30 output
  • Gemini 1.5 Flash: 0.075input/0.075 input / 0.30 output
  • GPT-4.1 Nano: 0.10input/0.10 input / 0.40 output
  • Gemini 2.5 Flash Lite: 0.10input/0.10 input / 0.40 output
  • GPT-4o Mini: 0.15input/0.15 input / 0.60 output
  • Claude 3 Haiku: 0.25input/0.25 input / 1.25 output
  • Claude 3.5 Haiku: 0.80input/0.80 input / 4.00 output
Perfect for: Classification, extraction, simple Q&A, high-throughput tasks
Balanced performance and cost:
  • Claude Haiku 4.5: 1.00input/1.00 input / 5.00 output
  • GPT-4.1: 2.00input/2.00 input / 8.00 output
  • GPT-4o: 2.50input/2.50 input / 10.00 output
  • Claude Sonnet 4.5: 3.00input/3.00 input / 15.00 output
  • Claude Opus 4.5: 5.00input/5.00 input / 25.00 output
Perfect for: Customer support, content generation, code assistance
Advanced reasoning and complex tasks:
  • Claude Opus 4: 15.00input/15.00 input / 75.00 output
  • o1: 15.00input/15.00 input / 60.00 output
  • o3: 2.00input/2.00 input / 8.00 output
  • o1-pro: 150.00input/150.00 input / 600.00 output
  • GPT-5 Pro: 15.00input/15.00 input / 120.00 output
  • GPT-4 (legacy): 30.00input/30.00 input / 60.00 output
Perfect for: Complex reasoning, research, code generation, expert analysis
AI image generation models:
  • Imagen 4 Fast: $0.02 per image
  • DALL-E 3 Standard (1024×1024): $0.040 per image
  • Imagen 4 Standard: $0.04 per image
  • gemini-2.5-flash-image: $0.039 per image (+ token costs)
  • Imagen 4 Ultra: $0.06 per image
  • DALL-E 3 HD (1024×1024): $0.080 per image
  • DALL-E 3 Large (1792×1024): 0.0800.080-0.120 per image
Perfect for: Marketing materials, product images, illustrations
AI video generation models:
  • Veo 3.0/3.1 Fast: $0.15 per second
  • Veo 3.0/3.1 Standard: $0.40 per second
Perfect for: Marketing videos, product demos, content creation
Audio transcription models:
  • Whisper-1: 0.006perminute(0.006 per minute (0.0001/second)
Text-to-speech models:
  • TTS-1 (Standard): $15.00 per 1M characters
  • TTS-1-HD (High Definition): $30.00 per 1M characters
Perfect for: Transcription services, voice assistants, audiobooks, accessibility

How Pricing Works

Token Calculation

LLM Ops tracks both input and output tokens separately:
  • Input tokens = Your prompt + any system messages + conversation history
  • Output tokens = The model’s response
Example:

Cost Calculation

Your total cost is calculated as:
All costs are tracked in real-time and displayed in your LLM Ops Dashboard.

Image & Video Generation Pricing

Image Generation Models

Models like DALL-E 3, Imagen 4, and gemini-2.5-flash-image generate images and are priced differently: DALL-E 3 (Per-Image):
Imagen 4 Models (Per-Image):
gemini-2.5-flash-image (Token-Based):

Video Generation Models

Veo Models (Per-Second):

Audio Transcription Models

Whisper (Per-Second):

Text-to-Speech Models

TTS Models (Per-Character):
Provider Counting:Image, video, and audio generation costs are calculated based on what the provider reports:
  • Images (DALL-E 3, Imagen 4): Provider returns number of images generated
  • Videos (Veo): Provider returns video duration in seconds
  • Audio (Whisper): Provider returns audio duration in seconds
  • TTS (TTS-1, TTS-1-HD): Provider returns character count of input text
We trust the provider’s counts and multiply by our pricing table.

Multimodal Pricing

Two Types of Multimodal Models

1. Multimodal Understanding Models (Token-Based)
  • These models analyze images, videos, and audio
  • Input media is converted to tokens by the provider
  • Charged per token (text + media tokens combined)
  • Examples: GPT-4o, Claude Opus, Gemini 2.5 Flash
2. Media Generation Models (Per-Unit)
  • These models create images, videos, or audio
  • DALL-E 3, Imagen 4: Charged per image generated
  • Veo 3: Charged per second of video generated
  • Whisper: Charged per second of audio transcribed
  • TTS: Charged per character of text input
  • gemini-2.5-flash-image: Hybrid (token-based, but image output uses premium rate)
How Multimodal Tokens Are Tracked:For understanding models, providers (OpenAI, Anthropic, Google) automatically convert images, video, and audio into tokens and include them in the response. LLM Ops tracks the total token count returned by the provider.Image/video/audio input tokens are included in input_tokens - they are not tracked separately.For generation models, we track based on what the provider charges:
  • Text tokens: Standard input/output pricing
  • Images generated: Per-image or per-token (depending on model)
  • Video generated: Per-second of video
  • Audio transcribed: Per-second of audio
  • TTS generated: Per-character of input text

Multimodal Understanding Models (Token-Based)

These models analyze images, video, and audio sent as input:
Models with image understanding:
  • Gemini 2.0/2.5 Flash
  • GPT-4o
  • GPT-4o Mini
  • Claude Opus 4/4.5
  • Claude Sonnet 4/4.5
How it works:
  1. You send an image with your prompt
  2. Provider converts image to tokens based on resolution
  3. Provider returns total input_tokens (text + image)
  4. LLM Ops tracks the total as input tokens
  5. Cost = input_tokens × input_price
Note: Higher resolution images = more input tokens = higher cost

Pricing Breakdown Summary

Here’s how different types of content are charged:
Key Takeaways:
  • Understanding (Input): Media → Tokens → Cost per token
  • Generation (Output):
    • Images: Per image (DALL-E 3, Imagen 4) OR per token (gemini-2.5-flash-image)
    • Video: Per second (Veo)
    • Audio Transcription: Per second (Whisper)
    • Text-to-Speech: Per character (OpenAI TTS) OR per token (Gemini TTS)
  • Provider Controls Conversion: We trust provider counts
  • Not Tracked Separately: Input media tokens are combined with text tokens in input_tokens

Example: Image Token Calculation

LLM Ops Dashboard shows:
  • Input Tokens: 769 (includes both text and image)
  • Output Tokens: 50
  • Total Cost: $0.0024225
Image/Audio/Video Breakdown Not Available:LLM Ops does not currently separate multimodal tokens from text tokens. All input tokens (text + image + video + audio) are tracked together as input_tokens.If you need separate multimodal token tracking, please contact us at support@cloudidr.com.

Pricing Updates

Model pricing is set by the providers (Anthropic, OpenAI, Google) and can change at any time. How we handle updates:
  • ✅ We monitor provider pricing pages daily
  • ✅ Updates are applied within 24 hours of provider changes
  • ✅ Historical data uses pricing from the time of request
  • ✅ You’re notified of major pricing changes
Last pricing update: January 10, 2026Check this page regularly for pricing updates.

Need a Model Added?

If you’re using a model that’s not listed here:
1

Check Provider Documentation

Verify the model exists in your provider’s official API docs
2

Contact Us

Email support@cloudidr.com with:
  • Model name
  • Provider (Anthropic/OpenAI/Google)
  • Link to provider pricing
3

We'll Add It

We typically add new models within 2-3 business days

Next Steps

Anthropic Integration

Start tracking Claude costs

OpenAI Integration

Start tracking GPT costs

Google Integration

Start tracking Gemini costs