Le Chonk Meets LiteLLM: Two 422s and a Runtime Patch
TLDR⌗
Mistral shipped a preview of Mistral Large 4 – officially nicknamed le Chonk – on the 6th of October, with open weights promised for the end of the month. I pointed my LiteLLM gateway at it, and every real client session died on the second request with a 422 from Mistral’s API. Two separate causes, both LiteLLM’s to fix, one of them with an upstream pull request that has been open since June:
- LiteLLM annotates assistant responses with its own
provider_specific_fields(MCP tool lists and results). Clients replay the assistant turn verbatim on the next request, and Mistral rejects the unknown field. - Mistral’s function-tool schema requires
parameters. Qwen Code deliberately omits it for tools that take no arguments, which OpenAI and vLLM are fine with. Mistral is not, and the error it returns points you at the wrong thing.
The fix is a sixty-line Python module that LiteLLM imports at startup through its own litellm_settings.callbacks hook, wrapping two upstream functions in-process. No image rebuild, nothing to redo on a version bump, and it fails loudly if upstream moves the furniture. The whole thing, including the Kubernetes wiring, is below.
Le Chonk, briefly⌗
Mistral announced Large 4 on October 6th as a public API preview: a natively multimodal mixture-of-experts model: 1.05 trillion parameters in total, 52 billion active per token, a 1.6B vision encoder and a one-million-token context window according to the model docs, which label it a public preview, and a promise that the weights will be released under an open licence by the end of October. The API identifier is mistral-large-4, metered at $1.36 per million input tokens and $4.18 per million output. VentureBeat and MarkTechPost have the usual write-ups; SiliconANGLE covers the roadmap bit. There is a placeholder on Hugging Face with a release ETA on it.
The benchmark story is self-reported at this stage, as it always is before weights ship, but the shape of it is interesting for anyone who runs their own models. Mistral is leaning hard on cybersecurity (93% of Cybench challenges, top five on the Artificial Analysis Cyber Index), claims a narrow win over GPT-6 Astra on visual grounding, and is candid that coding is middling: 61.7% on DeepSWE and 28.3% on Terminal-Bench 4, which is nowhere near the frontier. Artificial Analysis put the preview at 38 on their intelligence index, against 9 for Large 3. The Vals write-up is worth a read with the caveat that Mistral commissioned it.
Why I care, given that I run local models for a reason: if the weights do land in three weeks, a trillion-parameter MoE with fifty billion active is exactly the kind of thing that becomes runnable on serious home hardware in a quantised form, and a preview API is the cheapest way to find out whether it is worth the disk space. Which means plugging it into the same LiteLLM gateway everything else here goes through.
The setup, and the first 422⌗
My LiteLLM proxy is the one front door for every model on the network, local or cloud: LibreChat talks to it, Qwen Code talks to it, the CI review jobs talk to it, and it also runs LiteLLM’s MCP gateway, so clients get the same tool set whichever model they pick. Adding Mistral was one model_list entry:
- model_name: le-chonk
litellm_params:
model: mistral/mistral-large-4
api_base: https://api.mistral.ai/v1
api_key: os.environ/MISTRAL_API_KEY
First prompt: fine. Second prompt in any conversation that had used a tool: this.
{"detail":[{"type":"extra_forbidden",
"loc":["body","messages",1,"assistant","provider_specific_fields"],
"msg":"Extra inputs are not permitted",
"input":{"mcp_list_tools":[{"type":"function","function":{"name":"MetaMCP-ComfyUI__generate_image", ...
Mistral’s API validates request bodies with Pydantic in strict mode: any key it does not know, anywhere in the body, is a 422. Most providers ignore what they do not understand. Mistral refuses the lot.
The key it is refusing here is one LiteLLM itself put there. When a chat completion goes through LiteLLM’s MCP gateway, the proxy attaches provider_specific_fields to the assistant message in the response (or to the first streamed delta), carrying mcp_list_tools, mcp_tool_calls and mcp_call_results so a UI can show what happened. Any client that then replays the assistant turn verbatim – which is what the OpenAI SDK’s stream accumulators and most chat front-ends do – sends that field straight back on the next request, and LiteLLM’s Mistral transform does not strip it. It does strip reasoning_content and thinking_blocks, through a function called _strip_output_only_fields whose docstring explains exactly this failure mode, so somebody upstream got halfway there.
The rest of the way is BerriAI/litellm#30912, a fourteen-line pull request that drops provider_specific_fields, metadata and cache_control too. It was opened in June for issue #30882, and as I write this it is still open; the issue was closed by the stale bot in September. LiteLLM ships several releases a week, so I had assumed something this small would have been swept up. It had not, in 1.104.2, which came out this morning.
The second 422, which lies about itself⌗
With the first problem patched (how, in a minute), the LiteLLM playground got past its second turn. Then I pointed Qwen Code at it and got what looked like the same error back:
MistralException - {"detail":[
{"type":"literal_error","loc":["body","tools","list[union[WebSearchTool,...,Tool]]",7,"WebSearchTool","type"],
"msg":"Input should be <ConversationToolTypes.Websearch: 'web_search'>","input":"function"},
{"type":"extra_forbidden","loc":["body","tools","list[union[WebSearchTool,...,Tool]]",7,"WebSearchTool","function"],
"msg":"Extra inputs are not permitted","input":{"name":"get_goal","description":"Read the current Goal ...
Same phrase, different location: tools, item 7, a tool called get_goal. And the first two entries tell you nothing useful, because of how Pydantic reports a failed union. Mistral’s tools array accepts seven kinds of tool – web search, premium web search, code interpreter, image generation, document library, connector, and the plain function Tool – and when an item matches none of them, the error lists every variant’s complaint in turn. The first six are all variations on “this isn’t a web search tool”, which is true and irrelevant. The one that matters is the seventh, and it is at the bottom of a 72 KB error body.
Which you will not see. LiteLLM truncates the message to about 4 KB in its own log, truncates it again in the spend-logs table, and Qwen Code shows you one line of it. The way to read it is to stop trying: send a request with one tool and nothing else, so the body is small enough to print.
curl -s https://litellm.example/v1/chat/completions \
-H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
-d '{"model":"le-chonk","max_tokens":1,"messages":[{"role":"user","content":"hi"}],
"tools":[{"type":"function","function":{"name":"get_goal","description":"Read the current goal"}}]}'
Sixteen errors for one tool. Fifteen lines of union noise, and then:
tool[0] Tool.function.parameters: missing: Field required
Mistral’s Function model is name, parameters, description and strict, and parameters is not optional. OpenAI’s is optional: a function with no parameters simply takes no arguments, and vLLM, which serves my local models to the same client without complaint, agrees. Qwen Code’s OpenAI converter goes out of its way to omit the key: when a tool’s schema declares an empty argument list it sets parameters to undefined, which JSON serialisation drops entirely. Perfectly legal. Mistral says no, and then tells you about web search.
Add the same request with "parameters":{"type":"object","properties":{}} and it returns 200. That is the whole fix for this one: give argument-less tools an empty object schema, and while we are in there, drop any key that is not in Mistral’s four, because the next SDK will leak something else.
Patching LiteLLM without rebuilding it⌗
I have done the Dockerfile-with-a-sed thing before, for a one-line timeout fix in LibreChat, and I did not want to do it again: it means a build, a push to the registry, a tag to bump in two places, and all of it to redo on every upstream release. LiteLLM has a better door.
The proxy config’s litellm_settings.callbacks accepts a dotted path like my_module.my_instance, and the proxy resolves it relative to the directory containing config.yaml: it looks for my_module.py next to the config, loads it with importlib, and takes the named attribute, which has to be a CustomLogger. (The function is get_instance_fn in litellm/proxy/types_utils/utils.py, if you want to see for yourself.) The interesting part is not the callback. It is that importing the module runs its top-level code, once, inside the proxy process, before the first request. Anything that module does to LiteLLM’s own classes sticks for the life of the process.
So the patch is a module that wraps the two upstream functions:
"""
Runtime patches for the LiteLLM proxy.
The proxy imports this file at startup through `litellm_settings.callbacks` in
config.yaml (proxy_config in litellm-values.yaml). kustomize packs it into the
litellm-patches ConfigMap, which is mounted next to config.yaml as
/etc/litellm/litellm_patches.py, and LiteLLM resolves
`litellm_patches.proxy_handler_instance` relative to the config file's
directory. Module-level code therefore runs once per proxy process, before the
first request, inside the stock upstream image: no rebuild, and nothing to
re-push on a LiteLLM version bump. The ConfigMap name carries a content hash,
so editing this file and deploying rolls the pods.
Each patch probes that it took effect and raises otherwise. An exception here
fails config loading, so the new pod never becomes Ready and the rollout stalls
while the old pod keeps serving: a visible failure, not a silent regression.
=== Patch: Mistral rejects extra assistant-message fields =====================
Mistral's API validates request bodies strictly and answers 422
`extra_forbidden` for any unknown field in a message. LiteLLM attaches its own
bookkeeping to assistant responses - `provider_specific_fields`, which carries
`mcp_list_tools` for calls through the MCP gateway - and clients such as
LibreChat replay the assistant turn verbatim on the next request, so every
multi-turn tool conversation with a mistral/* model fails on the second call.
Upstream's MistralConfig._strip_output_only_fields drops only
`reasoning_content` and `thinking_blocks`; the fix that also drops
`provider_specific_fields`, `metadata` and `cache_control`
(BerriAI/litellm#30912, for issue #30882) has been open since 2026-06. This
wraps the upstream function to drop the same fields. It is harmless once
upstream does it too; remove it then.
=== Patch: Mistral's function-tool schema is strict and `parameters` is required
The same strict validation applies to `tools`. Mistral's Function model is
exactly `name`, `parameters` (required), `description` and `strict`, and the
tool object is `type` plus `function`. Two client habits break that:
- Qwen Code (a Gemini CLI fork) omits `parameters` entirely for a tool whose
schema declares no arguments, which OpenAI allows and Mistral rejects with
"Field required". Mistral reports the failure against every member of its
tool union, so the visible complaint is usually the WebSearchTool variant's
"Extra inputs are not permitted" rather than the real one.
- Anthropic-style clients add `cache_control` on the tool object, and other
SDKs leak extra keys inside `function`.
Upstream's MistralConfig._clean_tool_schema_for_mistral only strips
JSON-schema `$ref`s. This wraps it to drop unknown keys from function tools and
to default a missing `parameters` to an empty object schema. The schema's
contents are passed through untouched.
"""
from litellm.integrations.custom_logger import CustomLogger
from litellm.llms.mistral.chat.transformation import MistralConfig
MISTRAL_EXTRA_FIELDS = ("provider_specific_fields", "metadata", "cache_control")
def _patch_mistral_strip_extra_fields() -> None:
# AttributeError here means upstream moved or renamed the hook point:
# re-read litellm/llms/mistral/chat/transformation.py and adjust.
upstream = MistralConfig._strip_output_only_fields
if getattr(upstream, "_snowcoder_patch", None):
return # already applied (config reloads re-import this module)
def _strip(cls, message):
message = upstream(message)
if message.get("role") != "assistant":
return message
return {k: v for k, v in message.items() if k not in MISTRAL_EXTRA_FIELDS}
_strip._snowcoder_patch = "mistral-extra-fields"
MistralConfig._strip_output_only_fields = classmethod(_strip)
probe = {
"role": "assistant",
"content": "x",
"tool_calls": [],
"provider_specific_fields": {"mcp_list_tools": []},
"metadata": {},
"cache_control": {},
"reasoning_content": "r",
"thinking_blocks": [],
}
got = MistralConfig._strip_output_only_fields(dict(probe))
want = {"role": "assistant", "content": "x", "tool_calls": []}
if got != want:
raise RuntimeError(f"litellm_patches: Mistral strip patch did not take effect: {got!r}")
print(
"litellm_patches: Mistral assistant messages now also drop "
+ ", ".join(MISTRAL_EXTRA_FIELDS),
flush=True,
)
_patch_mistral_strip_extra_fields()
TOOL_KEYS = ("type", "function")
FUNCTION_KEYS = ("name", "description", "parameters", "strict")
NO_PARAMETERS = {"type": "object", "properties": {}}
def _patch_mistral_clean_tools() -> None:
upstream = MistralConfig._clean_tool_schema_for_mistral
if getattr(upstream, "_snowcoder_patch", None):
return
def _clean(cls, tools):
tools = upstream(tools)
if not tools:
return tools
cleaned = []
for tool in tools:
if isinstance(tool, dict) and tool.get("type") == "function" and isinstance(tool.get("function"), dict):
tool = {k: v for k, v in tool.items() if k in TOOL_KEYS}
fn = {k: v for k, v in tool["function"].items() if k in FUNCTION_KEYS}
if not isinstance(fn.get("parameters"), dict):
fn["parameters"] = dict(NO_PARAMETERS)
tool["function"] = fn
cleaned.append(tool)
return cleaned
_clean._snowcoder_patch = "mistral-tool-keys"
MistralConfig._clean_tool_schema_for_mistral = classmethod(_clean)
probe = [
{
"type": "function",
"function": {
"name": "f",
"description": "d",
"parameters": {"type": "object", "properties": {}},
"strict": False,
"parametersJsonSchema": {"type": "object"},
"response": {"type": "string"},
},
"cache_control": {"type": "ephemeral"},
},
{"type": "function", "function": {"name": "get_goal", "description": "parameterless"}},
{"type": "web_search"},
]
got = MistralConfig._clean_tool_schema_for_mistral(probe)
want = [
{
"type": "function",
"function": {"name": "f", "description": "d", "parameters": {"type": "object", "properties": {}}, "strict": False},
},
{"type": "function", "function": {"name": "get_goal", "description": "parameterless", "parameters": NO_PARAMETERS}},
{"type": "web_search"},
]
if got != want:
raise RuntimeError(f"litellm_patches: Mistral tool patch did not take effect: {got!r}")
print(
"litellm_patches: Mistral function tools reduced to " + ", ".join(FUNCTION_KEYS) + ", parameters defaulted when missing",
flush=True,
)
_patch_mistral_clean_tools()
class _NoOpLogger(CustomLogger):
"""litellm_settings.callbacks needs a CustomLogger instance to load this module; this one does nothing."""
proxy_handler_instance = _NoOpLogger()
Three design points, because they are the difference between a patch and a landmine:
Wrap, don’t replace. Each patch calls the upstream function first and only adds its own stripping afterwards. When #30912 or something like it finally merges, the wrapper becomes a no-op rather than a conflict, and nothing needs removing on the day.
Probe, and raise. After wrapping, each patch runs a tiny fixture through the patched function and raises if the result is not what it expects. An exception at import time fails the proxy’s config load, and the proxy exits. Under Kubernetes that means the new pod never passes readiness, the rollout stalls, and the old pod keeps serving – which is the failure mode you want from a thing that silently rewrites requests. A log line you did not read is not a safety net.
Only the one provider. Both wrappers live on MistralConfig, so cache_control still reaches Anthropic and thinking_blocks still round-trip where they should. Stripping these globally with a pre-call hook would have been shorter and wrong.
Wiring it in⌗
In my case the proxy runs from the upstream Helm chart, inflated by kustomize, so the module rides along as a content-hashed ConfigMap. kustomization.yaml:
configMapGenerator:
- name: litellm-patches
files:
- config/litellm_patches.py
generatorOptions:
disableNameSuffixHash: false # edit the file -> new name -> rollout
And in the chart values, mount it beside config.yaml (which the chart puts at /etc/litellm/config.yaml) and name it in the proxy config:
volumeMounts:
- name: litellm-patches
mountPath: /etc/litellm/litellm_patches.py
subPath: litellm_patches.py
readOnly: true
volumes:
- name: litellm-patches
configMap:
name: litellm-patches # kustomize rewrites the hashed name here
proxy_config:
litellm_settings:
callbacks:
- litellm_patches.proxy_handler_instance
If you run the proxy under docker-compose, the equivalent is a bind mount of the file next to your config and the same two lines of litellm_settings. There is nothing Kubernetes-specific in the mechanism.
One trap for the impatient: litellm --config config.yaml --skip_server_startup does not exercise the callbacks path, so it will exit cleanly whether or not your module even parses. To test locally, start the server properly and look for your own log lines:
litellm_patches: Mistral assistant messages now also drop provider_specific_fields, metadata, cache_control
litellm_patches: Mistral function tools reduced to name, description, parameters, strict, parameters defaulted when missing
INFO: Application startup complete.
And then test the negative case, because a safety net you have never seen catch anything is a rumour: add a raise to the module, start the server, and confirm you get Application startup failed. Exiting. rather than a running proxy.
Mileage⌗
After the rollout, the one-tool curl above returns 200 through the gateway with no parameters key at all; the playground’s second-turn 422 is gone; and a full Qwen Code session against le-chonk – a repository survey, half a dozen tool rounds with several calls each, every round replaying the previous assistant turn, argument-less tools in the set – ran to a tidy summary with no 422s. That is an afternoon’s testing, not a month’s, so treat it accordingly.
Two things I noticed and left alone. LiteLLM’s own playground, with MCP tools auto-executing, streams the tool round and then goes quiet instead of streaming the model’s follow-up answer; the spend logs show the follow-up call was never completed, which is a streaming bug in LiteLLM’s MCP path and nothing to do with Mistral. And Mistral’s strictness means any other client with its own habits – Anthropic-flavoured cache_control on tools, say – will surface as another 422 in another location, which the tool-key whitelist above should already cover, but I would not bet the house on the list being complete.
As for the model itself: early days, and this post is about the plumbing. Tool calling works once it can see the tools, and the question I actually care about – what it is like at home, quantised – waits on the weights.
Disclosure, in the spirit of transparency about AI use: this post was drafted by Claude, which also did the debugging it describes, from the one-tool reproduction to the patch module, with me driving the clients and the deploys; I edited it. The code is the code that is running. The cover comes from the ComfyUI deployment on the same cluster, as last time: Claude wrote the prompt and queued the job, and the cat’s opinion of the paperwork is entirely Z-Image Turbo’s.
