v1.65.4-stable

April 5, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

Deploy this version

Docker
Pip

docker run litellm
docker run
-e STORE_MODEL_IN_DB=True
-p 4000:4000
ghcr.io/berriai/litellm:main-v1.65.4-stable

pip install litellm

pip install litellm==1.65.4.post1

v1.65.0-stable - Model Context Protocol

March 30, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

v1.65.0-stable is live now. Here are the key highlights of this release:

MCP Support: Support for adding and using MCP servers on the LiteLLM proxy.
UI view total usage after 1M+ logs: You can now view usage analytics after crossing 1M+ logs in DB.

Model Context Protocol (MCP)

This release introduces support for centrally adding MCP servers on LiteLLM. This allows you to add MCP server endpoints and your developers can list and call MCP tools through LiteLLM.

UI view total usage after 1M+ logs

This release brings the ability to view total usage analytics even after exceeding 1M+ logs in your database. We've implemented a scalable architecture that stores only aggregate usage data, resulting in significantly more efficient queries and reduced database CPU utilization.

View total usage after 1M+ logs

How this works:
- We now aggregate usage data into a dedicated DailyUserSpend table, significantly reducing query load and CPU usage even beyond 1M+ logs.

Daily Spend Breakdown API:

Retrieve granular daily usage data (by model, provider, and API key) with a single endpoint. Example Request:

Daily Spend Breakdown API
curl -L -X GET 'http://localhost:4000/user/daily/activity?start_date=2025-03-20&end_date=2025-03-27' \
-H 'Authorization: Bearer sk-...'

Daily Spend Breakdown API Response
{
    "results": [
        {
            "date": "2025-03-27",
            "metrics": {
                "spend": 0.0177072,
                "prompt_tokens": 111,
                "completion_tokens": 1711,
                "total_tokens": 1822,
                "api_requests": 11
            },
            "breakdown": {
                "models": {
                    "gpt-4o-mini": {
                        "spend": 1.095e-05,
                        "prompt_tokens": 37,
                        "completion_tokens": 9,
                        "total_tokens": 46,
                        "api_requests": 1
                },
                "providers": { "openai": { ... }, "azure_ai": { ... } },
                "api_keys": { "3126b6eaf1...": { ... } }
            }
        }
    ],
    "metadata": {
        "total_spend": 0.7274667,
        "total_prompt_tokens": 280990,
        "total_completion_tokens": 376674,
        "total_api_requests": 14
    }
}

New Models / Updated Models

Support for Vertex AI gemini-2.0-flash-lite & Google AI Studio gemini-2.0-flash-lite PR
Support for Vertex AI Fine-Tuned LLMs PR
Nova Canvas image generation support PR
OpenAI gpt-4o-transcribe support PR
Added new Vertex AI text embedding model PR

LLM Translation

OpenAI Web Search Tool Call Support PR
Vertex AI topLogprobs support PR
Support for sending images and video to Vertex AI multimodal embedding Doc
Support litellm.api_base for Vertex AI + Gemini across completion, embedding, image_generation PR
Bug fix for returning response_cost when using litellm python SDK with LiteLLM Proxy PR
Support for max_completion_tokens on Mistral API PR
Refactored Vertex AI passthrough routes - fixes unpredictable behaviour with auto-setting default_vertex_region on router model add PR

Spend Tracking Improvements

Log 'api_base' on spend logs PR
Support for Gemini audio token cost tracking PR
Fixed OpenAI audio input token cost tracking PR

UI

Model Management

Allowed team admins to add/update/delete models on UI PR
Added render supports_web_search on model hub PR

Request Logs

Show API base and model ID on request logs PR
Allow viewing keyinfo on request logs PR

Usage Tab

Added Daily User Spend Aggregate view - allows UI Usage tab to work > 1m rows PR
Connected UI to "LiteLLM_DailyUserSpend" spend table PR

Logging Integrations

Fixed StandardLoggingPayload for GCS Pub Sub Logging Integration PR
Track litellm_model_name on StandardLoggingPayload Docs

Performance / Reliability Improvements

LiteLLM Redis semantic caching implementation PR
Gracefully handle exceptions when DB is having an outage PR
Allow Pods to startup + passing /health/readiness when allow_requests_on_db_unavailable: True and DB is down PR

General Improvements

Support for exposing MCP tools on litellm proxy PR
Support discovering Gemini, Anthropic, xAI models by calling their /v1/model endpoint PR
Fixed route check for non-proxy admins on JWT auth PR
Added baseline Prisma database migrations PR
View all wildcard models on /model/info PR

Security

Bumped next from 14.2.21 to 14.2.25 in UI dashboard PR

Complete Git Diff

Here's the complete git diff

v1.65.0 - Team Model Add - update

March 28, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

v1.65.0 updates the /model/new endpoint to prevent non-team admins from creating team models.

This means that only proxy admins or team admins can create team models.

Additional Changes

Allows team admins to call /model/update to update team models.
Allows team admins to call /model/delete to delete team models.
Introduces new user_models_only param to /v2/model/info - only return models added by this user.

These changes enable team admins to add and manage models for their team on the LiteLLM UI + API.

v1.63.14-stable

March 22, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

These are the changes since v1.63.11-stable.

This release brings:

LLM Translation Improvements (MCP Support and Bedrock Application Profiles)
Perf improvements for Usage-based Routing
Streaming guardrail support via websockets
Azure OpenAI client perf fix (from previous release)

Docker Run LiteLLM Proxy

docker run
-e STORE_MODEL_IN_DB=True
-p 4000:4000
ghcr.io/berriai/litellm:main-v1.63.14-stable.patch1

Demo Instance

Here's a Demo Instance to test changes:

Instance: https://demo.litellm.ai/
Login Credentials:
- Username: admin
- Password: sk-1234

New Models / Updated Models

Azure gpt-4o - fixed pricing to latest global pricing - PR
O1-Pro - add pricing + model information - PR
Azure AI - mistral 3.1 small pricing added - PR
Azure - gpt-4.5-preview pricing added - PR

LLM Translation

New LLM Features

Bedrock: Support bedrock application inference profiles Docs
- Infer aws region from bedrock application profile id - (arn:aws:bedrock:us-east-1:...)
Ollama - support calling via /v1/completions Get Started
Bedrock - support us.deepseek.r1-v1:0 model name Docs
OpenRouter - OPENROUTER_API_BASE env var support Docs
Azure - add audio model parameter support - Docs
OpenAI - PDF File support Docs
OpenAI - o1-pro Responses API streaming support Docs
[BETA] MCP - Use MCP Tools with LiteLLM SDK Docs

Bug Fixes

Voyage: prompt token on embedding tracking fix - PR
Sagemaker - Fix ‘Too little data for declared Content-Length’ error - PR
OpenAI-compatible models - fix issue when calling openai-compatible models w/ custom_llm_provider set - PR
VertexAI - Embedding ‘outputDimensionality’ support - PR
Anthropic - return consistent json response format on streaming/non-streaming - PR

Spend Tracking Improvements

litellm_proxy/ - support reading litellm response cost header from proxy, when using client sdk
Reset Budget Job - fix budget reset error on keys/teams/users PR
Streaming - Prevents final chunk w/ usage from being ignored (impacted bedrock streaming + cost tracking) PR

UI

Users Page
- Feature: Control default internal user settings PR
Icons:
- Feature: Replace external "artificialanalysis.ai" icons by local svg PR
Sign In/Sign Out
- Fix: Default login when default_user_id user does not exist in DB PR

Logging Integrations

Support post-call guardrails for streaming responses Get Started
Arize Get Started
- fix invalid package import PR
- migrate to using standardloggingpayload for metadata, ensures spans land successfully PR
- fix logging to just log the LLM I/O PR
- Dynamic API Key/Space param support Get Started
StandardLoggingPayload - Log litellm_model_name in payload. Allows knowing what the model sent to API provider was Get Started
Prompt Management - Allow building custom prompt management integration Get Started

Performance / Reliability improvements

Redis Caching - add 5s default timeout, prevents hanging redis connection from impacting llm calls PR
Allow disabling all spend updates / writes to DB - patch to allow disabling all spend updates to DB with a flag PR
Azure OpenAI - correctly re-use azure openai client, fixes perf issue from previous Stable release PR
Azure OpenAI - uses litellm.ssl_verify on Azure/OpenAI clients PR
Usage-based routing - Wildcard model support Get Started
Usage-based routing - Support batch writing increments to redis - reduces latency to same as ‘simple-shuffle’ PR
Router - show reason for model cooldown on ‘no healthy deployments available error’ PR
Caching - add max value limit to an item in in-memory cache (1MB) - prevents OOM errors on large image url’s being sent through proxy PR

General Improvements

Passthrough Endpoints - support returning api-base on pass-through endpoints Response Headers Docs
SSL - support reading ssl security level from env var - Allows user to specify lower security settings Get Started
Credentials - only poll Credentials table when STORE_MODEL_IN_DB is True PR
Image URL Handling - new architecture doc on image url handling Docs
OpenAI - bump to pip install "openai==1.68.2" PR
Gunicorn - security fix - bump gunicorn==23.0.0 PR

Complete Git Diff

Here's the complete git diff

v1.63.11-stable

March 15, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

These are the changes since v1.63.2-stable.

This release is primarily focused on:

[Beta] Responses API Support
Snowflake Cortex Support, Amazon Nova Image Generation
UI - Credential Management, re-use credentials when adding new models
UI - Test Connection to LLM Provider before adding a model

Known Issues

🚨 Known issue on Azure OpenAI - We don't recommend upgrading if you use Azure OpenAI. This version failed our Azure OpenAI load test

Docker Run LiteLLM Proxy

docker run
-e STORE_MODEL_IN_DB=True
-p 4000:4000
ghcr.io/berriai/litellm:main-v1.63.11-stable

Demo Instance

Here's a Demo Instance to test changes:

Instance: https://demo.litellm.ai/
Login Credentials:
- Username: admin
- Password: sk-1234

New Models / Updated Models

Image Generation support for Amazon Nova Canvas Getting Started
Add pricing for Jamba new models PR
Add pricing for Amazon EU models PR
Add Bedrock Deepseek R1 model pricing PR
Update Gemini pricing: Gemma 3, Flash 2 thinking update, LearnLM PR
Mark Cohere Embedding 3 models as Multimodal PR
Add Azure Data Zone pricing PR
- LiteLLM Tracks cost for azure/eu and azure/us models

LLM Translation

New Endpoints

[Beta] POST /responses API. Getting Started

New LLM Providers

Snowflake Cortex Getting Started

New LLM Features

Support OpenRouter reasoning_content on streaming Getting Started

Bug Fixes

OpenAI: Return code, param and type on bad request error More information on litellm exceptions
Bedrock: Fix converse chunk parsing to only return empty dict on tool use PR
Bedrock: Support extra_headers PR
Azure: Fix Function Calling Bug & Update Default API Version to 2025-02-01-preview PR
Azure: Fix AI services URL PR
Vertex AI: Handle HTTP 201 status code in response PR
Perplexity: Fix incorrect streaming response PR
Triton: Fix streaming completions bug PR
Deepgram: Support bytes.IO when handling audio files for transcription PR
Ollama: Fix "system" role has become unacceptable PR
All Providers (Streaming): Fix String data: stripped from entire content in streamed responses PR

Spend Tracking Improvements

Support Bedrock converse cache token tracking Getting Started
Cost Tracking for Responses API Getting Started
Fix Azure Whisper cost tracking Getting Started

UI

Re-Use Credentials on UI

You can now onboard LLM provider credentials on LiteLLM UI. Once these credentials are added you can re-use them when adding new models Getting Started

Test Connections before adding models

Before adding a model you can test the connection to the LLM provider to verify you have setup your API Base + API Key correctly

General UI Improvements

Add Models Page
- Allow adding Cerebras, Sambanova, Perplexity, Fireworks, Openrouter, TogetherAI Models, Text-Completion OpenAI on Admin UI
- Allow adding EU OpenAI models
- Fix: Instantly show edit + deletes to models
Keys Page
- Fix: Instantly show newly created keys on Admin UI (don't require refresh)
- Fix: Allow clicking into Top Keys when showing users Top API Key
- Fix: Allow Filter Keys by Team Alias, Key Alias and Org
- UI Improvements: Show 100 Keys Per Page, Use full height, increase width of key alias
Users Page
- Fix: Show correct count of internal user keys on Users Page
- Fix: Metadata not updating in Team UI
Logs Page
- UI Improvements: Keep expanded log in focus on LiteLLM UI
- UI Improvements: Minor improvements to logs page
- Fix: Allow internal user to query their own logs
- Allow switching off storing Error Logs in DB Getting Started
Sign In/Sign Out
- Fix: Correctly use PROXY_LOGOUT_URL when set Getting Started

Security

Support for Rotating Master Keys Getting Started
Fix: Internal User Viewer Permissions, don't allow internal_user_viewer role to see Test Key Page or Create Key Button More information on role based access controls
Emit audit logs on All user + model Create/Update/Delete endpoints Getting Started
JWT
- Support multiple JWT OIDC providers Getting Started
- Fix JWT access with Groups not working when team is assigned All Proxy Models access
Using K/V pairs in 1 AWS Secret Getting Started

Logging Integrations

Prometheus: Track Azure LLM API latency metric Getting Started
Athina: Added tags, user_feedback and model_options to additional_keys which can be sent to Athina Getting Started

Performance / Reliability improvements

Redis + litellm router - Fix Redis cluster mode for litellm router PR

General Improvements

OpenWebUI Integration - display thinking tokens

Guide on getting started with LiteLLM x OpenWebUI. Getting Started
Display thinking tokens on OpenWebUI (Bedrock, Anthropic, Deepseek) Getting Started

Complete Git Diff

Here's the complete git diff

v1.63.2-stable

March 8, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

These are the changes since v1.61.20-stable.

This release is primarily focused on:

LLM Translation improvements (more thinking content improvements)
UI improvements (Error logs now shown on UI)

info

This release will be live on 03/09/2025

Demo Instance

Here's a Demo Instance to test changes:

Instance: https://demo.litellm.ai/
Login Credentials:
- Username: admin
- Password: sk-1234

New Models / Updated Models

Add supports_pdf_input for specific Bedrock Claude models PR
Add pricing for amazon eu models PR
Fix Azure O1 mini pricing PR

LLM Translation

Support /openai/ passthrough for Assistant endpoints. Get Started
Bedrock Claude - fix tool calling transformation on invoke route. Get Started
Bedrock Claude - response_format support for claude on invoke route. Get Started
Bedrock - pass description if set in response_format. Get Started
Bedrock - Fix passing response_format: {"type": "text"}. PR
OpenAI - Handle sending image_url as str to openai. Get Started
Deepseek - return 'reasoning_content' missing on streaming. Get Started
Caching - Support caching on reasoning content. Get Started
Bedrock - handle thinking blocks in assistant message. Get Started
Anthropic - Return signature on streaming. Get Started

Note: We've also migrated from signature_delta to signature. Read more

Support format param for specifying image type. Get Started
Anthropic - /v1/messages endpoint - thinking param support. Get Started

Note: this refactors the [BETA] unified /v1/messages endpoint, to just work for the Anthropic API.

Vertex AI - handle $id in response schema when calling vertex ai. Get Started

Spend Tracking Improvements

Batches API - Fix cost calculation to run on retrieve_batch. Get Started
Batches API - Log batch models in spend logs / standard logging payload. Get Started

Management Endpoints / UI

Virtual Keys Page
- Allow team/org filters to be searchable on the Create Key Page
- Add created_by and updated_by fields to Keys table
- Show 'user_email' on key table
- Show 100 Keys Per Page, Use full height, increase width of key alias
Logs Page
- Show Error Logs on LiteLLM UI
- Allow Internal Users to View their own logs
Internal Users Page
- Allow admin to control default model access for internal users
Fix session handling with cookies

Logging / Guardrail Integrations

Fix prometheus metrics w/ custom metrics, when keys containing team_id make requests. PR

Performance / Loadbalancing / Reliability improvements

Cooldowns - Support cooldowns on models called with client side credentials. Get Started
Tag-based Routing - ensures tag-based routing across all endpoints (/embeddings, /image_generation, etc.). Get Started

General Proxy Improvements

Raise BadRequestError when unknown model passed in request
Enforce model access restrictions on Azure OpenAI proxy route
Reliability fix - Handle emoji’s in text - fix orjson error
Model Access Patch - don't overwrite litellm.anthropic_models when running auth checks
Enable setting timezone information in docker image

Complete Git Diff

Here's the complete git diff

v1.63.0 - Anthropic 'thinking' response update

March 5, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

v1.63.0 fixes Anthropic 'thinking' response on streaming to return the signature block. Github Issue

It also moves the response structure from signature_delta to signature to be the same as Anthropic. Anthropic Docs

Diff

"message": {
    ...
    "reasoning_content": "The capital of France is Paris.",
    "thinking_blocks": [
        {
            "type": "thinking",
            "thinking": "The capital of France is Paris.",
-            "signature_delta": "EqoBCkgIARABGAIiQL2UoU0b1OHYi+..." # 👈 OLD FORMAT
+            "signature": "EqoBCkgIARABGAIiQL2UoU0b1OHYi+..." # 👈 KEY CHANGE
        }
    ]
}

v1.61.20-stable

March 1, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

These are the changes since v1.61.13-stable.

This release is primarily focused on:

LLM Translation improvements (claude-3-7-sonnet + 'thinking'/'reasoning_content' support)
UI improvements (add model flow, user management, etc)

Demo Instance

Here's a Demo Instance to test changes:

Instance: https://demo.litellm.ai/
Login Credentials:
- Username: admin
- Password: sk-1234

New Models / Updated Models

Anthropic 3-7 sonnet support + cost tracking (Anthropic API + Bedrock + Vertex AI + OpenRouter)
1. Anthropic API Start here
2. Bedrock API Start here
3. Vertex AI API See here
4. OpenRouter See here
Gpt-4.5-preview support + cost tracking See here
Azure AI - Phi-4 cost tracking See here
Claude-3.5-sonnet - vision support updated on Anthropic API See here
Bedrock llama vision support See here
Cerebras llama3.3-70b pricing See here

LLM Translation

Infinity Rerank - support returning documents when return_documents=True Start here
Amazon Deepseek - <think> param extraction into ‘reasoning_content’ Start here
Amazon Titan Embeddings - filter out ‘aws_’ params from request body Start here
Anthropic ‘thinking’ + ‘reasoning_content’ translation support (Anthropic API, Bedrock, Vertex AI) Start here
VLLM - support ‘video_url’ Start here
Call proxy via litellm SDK: Support litellm_proxy/ for embedding, image_generation, transcription, speech, rerank Start here
OpenAI Pass-through - allow using Assistants GET, DELETE on /openai pass through routes Start here
Message Translation - fix openai message for assistant msg if role is missing - openai allows this
O1/O3 - support ‘drop_params’ for o3-mini and o1 parallel_tool_calls param (not supported currently) See here

Spend Tracking Improvements

Cost tracking for rerank via Bedrock See PR
Anthropic pass-through - fix race condition causing cost to not be tracked See PR
Anthropic pass-through: Ensure accurate token counting See PR

Management Endpoints / UI

Models Page - Allow sorting models by ‘created at’
Models Page - Edit Model Flow Improvements
Models Page - Fix Adding Azure, Azure AI Studio models on UI
Internal Users Page - Allow Bulk Adding Internal Users on UI
Internal Users Page - Allow sorting users by ‘created at’
Virtual Keys Page - Allow searching for UserIDs on the dropdown when assigning a user to a team See PR
Virtual Keys Page - allow creating a user when assigning keys to users See PR
Model Hub Page - fix text overflow issue See PR
Admin Settings Page - Allow adding MSFT SSO on UI
Backend - don't allow creating duplicate internal users in DB

Helm

support ttlSecondsAfterFinished on the migration job - See PR
enhance migrations job with additional configurable properties - See PR

Logging / Guardrail Integrations

Arize Phoenix support
‘No-log’ - fix ‘no-log’ param support on embedding calls

Performance / Loadbalancing / Reliability improvements

Single Deployment Cooldown logic - Use allowed_fails or allowed_fail_policy if set Start here

General Proxy Improvements

Hypercorn - fix reading / parsing request body
Windows - fix running proxy in windows
DD-Trace - fix dd-trace enablement on proxy

Complete Git Diff

View the complete git diff here.

v1.59.8-stable

January 31, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

info

Get a 7 day free trial for LiteLLM Enterprise here.

no call needed

New Models / Updated Models

New OpenAI /image/variations endpoint BETA support Docs
Topaz API support on OpenAI /image/variations BETA endpoint Docs
Deepseek - r1 support w/ reasoning_content (Deepseek API, Vertex AI, Bedrock)
Azure - Add azure o1 pricing See Here
Anthropic - handle -latest tag in model for cost calculation
Gemini-2.0-flash-thinking - add model pricing (it’s 0.0) See Here
Bedrock - add stability sd3 model pricing See Here (s/o Marty Sullivan)
Bedrock - add us.amazon.nova-lite-v1:0 to model cost map See Here
TogetherAI - add new together_ai llama3.3 models See Here

LLM Translation

LM Studio -> fix async embedding call
Gpt 4o models - fix response_format translation
Bedrock nova - expand supported document types to include .md, .csv, etc. Start Here
Bedrock - docs on IAM role based access for bedrock - Start Here
Bedrock - cache IAM role credentials when used
Google AI Studio (gemini/) - support gemini 'frequency_penalty' and 'presence_penalty'
Azure O1 - fix model name check
WatsonX - ZenAPIKey support for WatsonX Docs
Ollama Chat - support json schema response format Start Here
Bedrock - return correct bedrock status code and error message if error during streaming
Anthropic - Supported nested json schema on anthropic calls
OpenAI - metadata param preview support
1. SDK - enable via litellm.enable_preview_features = True
2. PROXY - enable via litellm_settings::enable_preview_features: true
Replicate - retry completion response on status=processing

Spend Tracking Improvements

Bedrock - QA asserts all bedrock regional models have same supported_ as base model
Bedrock - fix bedrock converse cost tracking w/ region name specified
Spend Logs reliability fix - when user passed in request body is int instead of string
Ensure ‘base_model’ cost tracking works across all endpoints
Fixes for Image generation cost tracking
Anthropic - fix anthropic end user cost tracking
JWT / OIDC Auth - add end user id tracking from jwt auth

Management Endpoints / UI

allows team member to become admin post-add (ui + endpoints)
New edit/delete button for updating team membership on UI
If team admin - show all team keys
Model Hub - clarify cost of models is per 1m tokens
Invitation Links - fix invalid url generated
New - SpendLogs Table Viewer - allows proxy admin to view spend logs on UI
1. New spend logs - allow proxy admin to ‘opt in’ to logging request/response in spend logs table - enables easier abuse detection
2. Show country of origin in spend logs
3. Add pagination + filtering by key name/team name
/key/delete - allow team admin to delete team keys
Internal User ‘view’ - fix spend calculation when team selected
Model Analytics is now on Free
Usage page - shows days when spend = 0, and round spend on charts to 2 sig figs
Public Teams - allow admins to expose teams for new users to ‘join’ on UI - Start Here
Guardrails
1. set/edit guardrails on a virtual key
2. Allow setting guardrails on a team
3. Set guardrails on team create + edit page
Support temporary budget increases on /key/update - new temp_budget_increase and temp_budget_expiry fields - Start Here
Support writing new key alias to AWS Secret Manager - on key rotation Start Here

Helm

add securityContext and pull policy values to migration job (s/o https://github.com/Hexoplon)
allow specifying envVars on values.yaml
new helm lint test

Logging / Guardrail Integrations

Log the used prompt when prompt management used. Start Here
Support s3 logging with team alias prefixes - Start Here
Prometheus Start Here
1. fix litellm_llm_api_time_to_first_token_metric not populating for bedrock models
2. emit remaining team budget metric on regular basis (even when call isn’t made) - allows for more stable metrics on Grafana/etc.
3. add key and team level budget metrics
4. emit litellm_overhead_latency_metric
5. Emit litellm_team_budget_reset_at_metric and litellm_api_key_budget_remaining_hours_metric
Datadog - support logging spend tags to Datadog. Start Here
Langfuse - fix logging request tags, read from standard logging payload
GCS - don’t truncate payload on logging
New GCS Pub/Sub logging support Start Here
Add AIM Guardrails support Start Here

Security

New Enterprise SLA for patching security vulnerabilities. See Here
Hashicorp - support using vault namespace for TLS auth. Start Here
Azure - DefaultAzureCredential support

Health Checks

Cleanup pricing-only model names from wildcard route list - prevent bad health checks
Allow specifying a health check model for wildcard routes - https://docs.litellm.ai/docs/proxy/health#wildcard-routes
New ‘health_check_timeout ‘ param with default 1min upperbound to prevent bad model from health check to hang and cause pod restarts. Start Here
Datadog - add data dog service health check + expose new /health/services endpoint. Start Here

Performance / Reliability improvements

3x increase in RPS - moving to orjson for reading request body
LLM Routing speedup - using cached get model group info
SDK speedup - using cached get model info helper - reduces CPU work to get model info
Proxy speedup - only read request body 1 time per request
Infinite loop detection scripts added to codebase
Bedrock - pure async image transformation requests
Cooldowns - single deployment model group if 100% calls fail in high traffic - prevents an o1 outage from impacting other calls
Response Headers - return
1. x-litellm-timeout
2. x-litellm-attempted-retries
3. x-litellm-overhead-duration-ms
4. x-litellm-response-duration-ms
ensure duplicate callbacks are not added to proxy
Requirements.txt - bump certifi version

General Proxy Improvements

JWT / OIDC Auth - new enforce_rbac param,allows proxy admin to prevent any unmapped yet authenticated jwt tokens from calling proxy. Start Here
fix custom openapi schema generation for customized swagger’s
Request Headers - support reading x-litellm-timeout param from request headers. Enables model timeout control when using Vercel’s AI SDK + LiteLLM Proxy. Start Here
JWT / OIDC Auth - new role based permissions for model authentication. See Here

Complete Git Diff

This is the diff between v1.57.8-stable and v1.59.8-stable.

Use this to see the changes in the codebase.

Git Diff

v1.59.0

January 17, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

info

Get a 7 day free trial for LiteLLM Enterprise here.

no call needed

UI Improvements

[Opt In] Admin UI - view messages / responses

You can now view messages and response logs on Admin UI.

How to enable it - add store_prompts_in_spend_logs: true to your proxy_config.yaml

Once this flag is enabled, your messages and responses will be stored in the LiteLLM_Spend_Logs table.

general_settings:
  store_prompts_in_spend_logs: true

DB Schema Change

Added messages and responses to the LiteLLM_Spend_Logs table.

By default this is not logged. If you want messages and responses to be logged, you need to opt in with this setting

general_settings:
  store_prompts_in_spend_logs: true

v1.57.8-stable

January 11, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

alerting, prometheus, secret management, management endpoints, ui, prompt management, finetuning, batch

New / Updated Models

Mistral large pricing - https://github.com/BerriAI/litellm/pull/7452
Cohere command-r7b-12-2024 pricing - https://github.com/BerriAI/litellm/pull/7553/files
Voyage - new models, prices and context window information - https://github.com/BerriAI/litellm/pull/7472
Anthropic - bump Bedrock claude-3-5-haiku max_output_tokens to 8192

General Proxy Improvements

Health check support for realtime models
Support calling Azure realtime routes via virtual keys
Support custom tokenizer on /utils/token_counter - useful when checking token count for self-hosted models
Request Prioritization - support on /v1/completion endpoint as well

LLM Translation Improvements

Deepgram STT support. Start Here
OpenAI Moderations - omni-moderation-latest support. Start Here
Azure O1 - fake streaming support. This ensures if a stream=true is passed, the response is streamed. Start Here
Anthropic - non-whitespace char stop sequence handling - PR
Azure OpenAI - support entrata id username + password based auth. Start Here
LM Studio - embedding route support. Start Here
WatsonX - ZenAPIKeyAuth support. Start Here

Prompt Management Improvements

Langfuse integration
HumanLoop integration
Support for using load balanced models
Support for loading optional params from prompt manager

Start Here

Finetuning + Batch APIs Improvements

Improved unified endpoint support for Vertex AI finetuning - PR
Add support for retrieving vertex api batch jobs - PR

NEW Alerting Integration

PagerDuty Alerting Integration.

Handles two types of alerts:

High LLM API Failure Rate. Configure X fails in Y seconds to trigger an alert.
High Number of Hanging LLM Requests. Configure X hangs in Y seconds to trigger an alert.

Start Here

Prometheus Improvements

Added support for tracking latency/spend/tokens based on custom metrics. Start Here

NEW Hashicorp Secret Manager Support

Support for reading credentials + writing LLM API keys. Start Here

Management Endpoints / UI Improvements

Create and view organizations + assign org admins on the Proxy UI
Support deleting keys by key_alias
Allow assigning teams to org on UI
Disable using ui session token for 'test key' pane
Show model used in 'test key' pane
Support markdown output in 'test key' pane

Helm Improvements

Prevent istio injection for db migrations cron job
allow using migrationJob.enabled variable within job

Logging Improvements

braintrust logging: respect project_id, add more metrics - https://github.com/BerriAI/litellm/pull/7613
Athina - support base url - ATHINA_BASE_URL
Lunary - Allow passing custom parent run id to LLM Calls

Git Diff

This is the diff between v1.56.3-stable and v1.57.8-stable.

Use this to see the changes in the codebase.

Git Diff

v1.57.7

January 10, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

langfuse, management endpoints, ui, prometheus, secret management

Langfuse Prompt Management

Langfuse Prompt Management is being labelled as BETA. This allows us to iterate quickly on the feedback we're receiving, and making the status clearer to users. We expect to make this feature to be stable by next month (February 2025).

Changes:

Include the client message in the LLM API Request. (Previously only the prompt template was sent, and the client message was ignored).
Log the prompt template in the logged request (e.g. to s3/langfuse).
Log the 'prompt_id' and 'prompt_variables' in the logged request (e.g. to s3/langfuse).

Start Here

Team/Organization Management + UI Improvements

Managing teams and organizations on the UI is now easier.

Changes:

Support for editing user role within team on UI.
Support updating team member role to admin via api - /team/member_update
Show team admins all keys for their team.
Add organizations with budgets
Assign teams to orgs on the UI
Auto-assign SSO users to teams

Start Here

Hashicorp Vault Support

We now support writing LiteLLM Virtual API keys to Hashicorp Vault.

Start Here

Custom Prometheus Metrics

Define custom prometheus metrics, and track usage/latency/no. of requests against them

This allows for more fine-grained tracking - e.g. on prompt template passed in request metadata

Start Here

v1.57.3 - New Base Docker Image

January 8, 2025

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

docker image, security, vulnerability

0 Critical/High Vulnerabilities

What changed?

LiteLLMBase image now uses cgr.dev/chainguard/python:latest-dev

Why the change?

To ensure there are 0 critical/high vulnerabilities on LiteLLM Docker Image

Migration Guide

If you use a custom dockerfile with litellm as a base image + apt-get

Instead of apt-get use apk, the base litellm image will no longer have apt-get installed.

You are only impacted if you use apt-get in your Dockerfile

# Use the provided base image
FROM ghcr.io/berriai/litellm:main-latest

# Set the working directory
WORKDIR /app

# Install dependencies - CHANGE THIS to `apk`
RUN apt-get update && apt-get install -y dumb-init 

Before Change

RUN apt-get update && apt-get install -y dumb-init

After Change

RUN apk update && apk add --no-cache dumb-init

v1.56.4

December 29, 2024

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

deepgram, fireworks ai, vision, admin ui, dependency upgrades

New Models

Deepgram Speech to Text

New Speech to Text support for Deepgram models. Start Here

from litellm import transcription
import os 

# set api keys 
os.environ["DEEPGRAM_API_KEY"] = ""
audio_file = open("/path/to/audio.mp3", "rb")

response = transcription(model="deepgram/nova-2", file=audio_file)

print(f"response: {response}")

Fireworks AI - Vision support for all models

LiteLLM supports document inlining for Fireworks AI models. This is useful for models that are not vision models, but still need to parse documents/images/etc. LiteLLM will add #transform=inline to the url of the image_url, if the model is not a vision model See Code

Proxy Admin UI

Test Key Tab displays model used in response

Test Key Tab renders content in .md, .py (any code/markdown format)

Dependency Upgrades

(Security fix) Upgrade to fastapi==0.115.5 https://github.com/BerriAI/litellm/pull/7447

Bug Fixes

Add health check support for realtime models Here
Health check error with audio_transcription model https://github.com/BerriAI/litellm/issues/5999

v1.56.3

December 28, 2024

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

guardrails, logging, virtual key management, new models

info

Get a 7 day free trial for LiteLLM Enterprise here.

no call needed

New Features

✨ Log Guardrail Traces

Track guardrail failure rate and if a guardrail is going rogue and failing requests. Start here

Traced Guardrail Success

Traced Guardrail Failure

`/guardrails/list`

/guardrails/list allows clients to view available guardrails + supported guardrail params

curl -X GET 'http://0.0.0.0:4000/guardrails/list'

Expected response

{
    "guardrails": [
        {
        "guardrail_name": "aporia-post-guard",
        "guardrail_info": {
            "params": [
            {
                "name": "toxicity_score",
                "type": "float",
                "description": "Score between 0-1 indicating content toxicity level"
            },
            {
                "name": "pii_detection",
                "type": "boolean"
            }
            ]
        }
        }
    ]
}

✨ Guardrails with Mock LLM

Send mock_response to test guardrails without making an LLM call. More info on mock_response here

curl -i http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-npnwjPQciVRok5yNZgKmFQ" \
  -d '{
    "model": "gpt-3.5-turbo",
    "messages": [
      {"role": "user", "content": "hi my email is ishaan@berri.ai"}
    ],
    "mock_response": "This is a mock response",
    "guardrails": ["aporia-pre-guard", "aporia-post-guard"]
  }'

Assign Keys to Users

You can now assign keys to users via Proxy UI

New Models

openrouter/openai/o1
vertex_ai/mistral-large@2411

Fixes

Fix vertex_ai/ mistral model pricing: https://github.com/BerriAI/litellm/pull/7345
Missing model_group field in logs for aspeech call types https://github.com/BerriAI/litellm/pull/7392

v1.56.1

December 27, 2024

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

key management, budgets/rate limits, logging, guardrails

info

Get a 7 day free trial for LiteLLM Enterprise here.

no call needed

✨ Budget / Rate Limit Tiers

Define tiers with rate limits. Assign them to keys.

Use this to control access and budgets across a lot of keys.

Start here

curl -L -X POST 'http://0.0.0.0:4000/budget/new' \
-H 'Authorization: Bearer sk-1234' \
-H 'Content-Type: application/json' \
-d '{
    "budget_id": "high-usage-tier",
    "model_max_budget": {
        "gpt-4o": {"rpm_limit": 1000000}
    }
}'

OTEL Bug Fix

LiteLLM was double logging litellm_request span. This is now fixed.

Relevant PR

Logging for Finetuning Endpoints

Logs for finetuning requests are now available on all logging providers (e.g. Datadog).

What's logged per request:

file_id
finetuning_job_id
any key/team metadata

Start Here:

Dynamic Params for Guardrails

You can now set custom parameters (like success threshold) for your guardrails in each request.

See guardrails spec for more details

v1.55.10

December 24, 2024

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

batches, guardrails, team management, custom auth

info

Get a free 7-day LiteLLM Enterprise trial here. Start here

No call needed

✨ Cost Tracking, Logging for Batches API (`/batches`)

Track cost, usage for Batch Creation Jobs. Start here

✨ `/guardrails/list` endpoint

Show available guardrails to users. Start here

✨ Allow teams to add models

This enables team admins to call their own finetuned models via litellm proxy. Start here

✨ Common checks for custom auth

Calling the internal common_checks function in custom auth is now enforced as an enterprise feature. This allows admins to use litellm's default budget/auth checks within their custom auth implementation. Start here

✨ Assigning team admins

Team admins is graduating from beta and moving to our enterprise tier. This allows proxy admins to allow others to manage keys/models for their own teams (useful for projects in production). Start here

v1.55.8-stable

December 22, 2024

Krrish Dholakia

CEO, LiteLLM

Ishaan Jaffer

CTO, LiteLLM

A new LiteLLM Stable release just went out. Here are 5 updates since v1.52.2-stable.

langfuse, fallbacks, new models, azure_storage

Langfuse Prompt Management

This makes it easy to run experiments or change the specific models gpt-4o to gpt-4o-mini on Langfuse, instead of making changes in your applications. Start here

Control fallback prompts client-side

Claude prompts are different than OpenAI

Pass in prompts specific to model when doing fallbacks. Start here

New Providers / Models

NVIDIA Triton /infer endpoint. Start here
Infinity Rerank Models Start here

✨ Azure Data Lake Storage Support

Send LLM usage (spend, tokens) data to Azure Data Lake. This makes it easy to consume usage data on other services (eg. Databricks) Start here

Docker Run LiteLLM

docker run \
-e STORE_MODEL_IN_DB=True \
-p 4000:4000 \
ghcr.io/berriai/litellm:litellm_stable_release_branch-v1.55.8-stable

Get Daily Updates

LiteLLM ships new releases every day. Follow us on LinkedIn to get daily updates.

Deploy this version​

Model Context Protocol (MCP)​

UI view total usage after 1M+ logs​

New Models / Updated Models​

LLM Translation​

Spend Tracking Improvements​

UI​

Model Management​

Request Logs​

Usage Tab​

Logging Integrations​

Performance / Reliability Improvements​

General Improvements​

Security​

Complete Git Diff​

Additional Changes​

Docker Run LiteLLM Proxy​

Demo Instance​

New Models / Updated Models​

LLM Translation​

Spend Tracking Improvements​

UI​

Logging Integrations​

Performance / Reliability improvements​

General Improvements​

Complete Git Diff​

Known Issues​

Docker Run LiteLLM Proxy​

Demo Instance​

New Models / Updated Models​

LLM Translation​

Spend Tracking Improvements​

UI​

Re-Use Credentials on UI​

Test Connections before adding models​

General UI Improvements​

Security​

Logging Integrations​

Performance / Reliability improvements​

General Improvements​

Complete Git Diff​

Demo Instance​

New Models / Updated Models​

LLM Translation​

Spend Tracking Improvements​

Management Endpoints / UI​

Logging / Guardrail Integrations​

Performance / Loadbalancing / Reliability improvements​

General Proxy Improvements​

Complete Git Diff​

Diff​

Demo Instance​

New Models / Updated Models​

LLM Translation​

Spend Tracking Improvements​

Management Endpoints / UI​

Helm​

Logging / Guardrail Integrations​

Performance / Loadbalancing / Reliability improvements​

General Proxy Improvements​

Complete Git Diff​

New Models / Updated Models​

LLM Translation​

Spend Tracking Improvements​

Management Endpoints / UI​

Helm​

Logging / Guardrail Integrations​

Security​

Health Checks​

Performance / Reliability improvements​

General Proxy Improvements​

Complete Git Diff​

UI Improvements​

[Opt In] Admin UI - view messages / responses​

DB Schema Change​

New / Updated Models​

General Proxy Improvements​

LLM Translation Improvements​

Prompt Management Improvements​

Finetuning + Batch APIs Improvements​

Deploy this version

Model Context Protocol (MCP)

UI view total usage after 1M+ logs

New Models / Updated Models

LLM Translation

Spend Tracking Improvements

UI

Model Management

Request Logs

Usage Tab

Logging Integrations

Performance / Reliability Improvements

General Improvements

Security

Complete Git Diff

Additional Changes

Docker Run LiteLLM Proxy

Demo Instance

New Models / Updated Models

LLM Translation

Spend Tracking Improvements

UI

Logging Integrations

Performance / Reliability improvements

General Improvements

Complete Git Diff

Known Issues

Docker Run LiteLLM Proxy

Demo Instance

New Models / Updated Models

LLM Translation

Spend Tracking Improvements

UI

Re-Use Credentials on UI

Test Connections before adding models

General UI Improvements

Security

Logging Integrations

Performance / Reliability improvements

General Improvements

Complete Git Diff

Demo Instance

New Models / Updated Models

LLM Translation

Spend Tracking Improvements

Management Endpoints / UI

Logging / Guardrail Integrations

Performance / Loadbalancing / Reliability improvements

General Proxy Improvements

Complete Git Diff

Diff

Demo Instance

New Models / Updated Models

LLM Translation

Spend Tracking Improvements

Management Endpoints / UI

Helm

Logging / Guardrail Integrations

Performance / Loadbalancing / Reliability improvements

General Proxy Improvements

Complete Git Diff

New Models / Updated Models

LLM Translation

Spend Tracking Improvements

Management Endpoints / UI

Helm

Logging / Guardrail Integrations

Security

Health Checks

Performance / Reliability improvements

General Proxy Improvements

Complete Git Diff

UI Improvements

[Opt In] Admin UI - view messages / responses

DB Schema Change

New / Updated Models

General Proxy Improvements

LLM Translation Improvements

Prompt Management Improvements

Finetuning + Batch APIs Improvements