TLDR⌗

Yesterday’s post got Mistral Large 4 – le Chonk – working through LiteLLM by patching two things Mistral’s API refuses. This one is about a thing Mistral’s API accepts perfectly happily, and never receives.

  • Large 4 is a reasoning model and reasons by default. Mistral’s documentation says to replay the full assistant message, thinking included, on every subsequent turn, and warns that stripping the thinking “significantly degrades output quality”.
  • LiteLLM hands the thinking to your client as reasoning_content, the client dutifully sends it back, and LiteLLM’s Mistral transform deletes it before the request leaves the building. A second code path flattens the content list, so even a client that sends Mistral’s own chunk format next to text loses it.
  • The fix is a third function in the same runtime patch module as before: rebuild Mistral’s thinking chunk from reasoning_content and protect it from the flattener. There is a thirty-second test, involving a pelican, that tells you whether any gateway and model combination is actually seeing replayed reasoning.
  • The difference in a real coding-agent session was not subtle.

Where this came from⌗

Having got the 422s out of the way, I was reading the patch module and noticed that the upstream function we wrap exists in the first place to drop reasoning_content and thinking_blocks from assistant messages before they go to Mistral. That seemed odd. The current generation of models – Qwen3’s recent releases, DeepSeek in thinking mode, Kimi K2, Claude, Gemini – all want their reasoning preserved across turns of a tool-calling loop, and I had gone to some trouble to set my local llama.cpp instance up to do exactly that. Was LiteLLM quietly undoing it?

The answer turned out to be: not for my local models, yes for Mistral, and Mistral is the one that minds most.

What LiteLLM does with reasoning, per provider⌗

I had the LiteLLM 1.104.2 source to hand from the first round, so this is from reading the code rather than the docs. On the way out LiteLLM is consistent: whatever a provider’s native thinking format is, it becomes reasoning_content on the OpenAI-shaped response, and in streaming it arrives as reasoning_content deltas. On the way back in, when a client replays the assistant turn, it does one of four things depending on the provider:

  • Strips it. Mistral, hosted vLLM and Databricks delete reasoning_content and thinking_blocks from assistant messages; Fireworks drops thinking_blocks. Together strips thinking_blocks but deliberately keeps reasoning_content, with a comment noting that Together only consumes it when the chat template is told not to clear thinking.

  • Translates it into the provider’s native format. Anthropic turns thinking_blocks back into signed thinking content blocks, dropping any it cannot sign. Bedrock Converse does the same into its reasoningContent blocks. Gemini replays reasoning_content as a thought part and signed blocks as thought signatures. Ollama maps it to its thinking field. This is the careful path, and it exists where the provider either signs its reasoning or errors without it.

  • Requires it, and fakes it if missing. DeepSeek in thinking mode and Moonshot’s reasoning models insist on reasoning_content on prior assistant turns. If the client didn’t send it, LiteLLM promotes it from provider_specific_fields or injects a single space and logs a warning that quality may suffer.

  • Leaves it alone. The plain OpenAI transform, and therefore every endpoint registered with the openai/ prefix, has no request-side handling of reasoning at all. Whatever the client sends goes through verbatim and it is the backend’s problem.

So “LiteLLM strips thinking” is neither true nor false; it is a per-provider decision, and for Mistral the decision is to strip. That decision arrived in June’s staging bundle (BerriAI/litellm#30968) and is unchanged on main. The reason is understandable: Mistral’s strict validation 422s on an unknown reasoning_content key, and deleting the key is the smallest change that makes the 422 go away. It is just not what Mistral asked for.

What Mistral asked for⌗

Mistral’s reasoning documentation lists mistral-large-4-0 among the models that take reasoning_effort (the values are high and none), and describes the response shape: with reasoning on, message.content is a list of chunks rather than a string, a ThinkChunk of type thinking holding the trace and a TextChunk holding the answer. Then, on history:

In multi-turn conversations, always replay the full assistant message (including ThinkChunk) back into the message history. Do not strip ThinkChunk from assistant messages before replaying them.

followed by the observation that stripping improves token efficiency but significantly degrades output quality.

Two details matter for what follows. First, Large 4 reasons whether or not you ask it to: a bare request through my gateway with no reasoning_effort came back with a reasoning trace attached. This is not an opt-in feature you can avoid by not opting in. Second, the thing to replay is a content chunk, not a top-level field. The OpenAI-style reasoning_content key that every client sends is the wrong shape for Mistral regardless of whether LiteLLM strips it; somebody has to translate.

What the clients do⌗

Somebody, it turns out, is Mistral’s own coding agent. Mistral Vibe keeps reasoning_content on its internal message type like everyone else, and its Mistral backend rebuilds a ThinkChunk from it on every assistant turn whenever the model’s thinking level is anything other than off. Its thinking levels map onto reasoning_effort: low becomes none, medium and above become high. The Vercel AI SDK fixed the same class of bug in @ai-sdk/mistral 3.0.52 by mapping reasoning parts to thinking content parts on replay. So the reference behaviour exists, and it is what Mistral’s own tooling does.

Qwen Code, which is what I actually drive Le Chonk with, has two paths and neither of them gets there. Its generic OpenAI-compatible provider sends reasoning_content back as a top-level key, which is correct for vLLM and llama.cpp and is what my gateway receives. It also has a dedicated Mistral provider, selected when the base URL is api.mistral.ai or the model name contains mistral, magistral, devstral and so on, and that provider’s entire job is to delete reasoning_content so Mistral doesn’t 422. Through a gateway under the name le-chonk you get the first path; talking to Mistral directly you get the second. Either way the model never sees its earlier reasoning.

I want to be fair to Qwen Code here: deleting the field is a perfectly reasonable response to a 422, and until a few weeks ago Mistral’s reasoning models were a niche. But the net effect for anyone running Large 4 behind LiteLLM was that two independent layers were each solving the 422 by throwing the thinking away, and the one layer that could have translated it was the one in the middle.

The amnesia test⌗

Because I cannot see the body LiteLLM sends to Mistral, I wanted a test that proves the thinking arrives, not just that the request is accepted. The trick is to put information in the replayed reasoning that the model could only know if it read it. Here it is as a curl against the proxy, with a tool so that it exercises the tool-loop shape:

curl -s "$LITELLM/v1/chat/completions" \
  -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' -d '{
  "model": "le-chonk", "max_tokens": 400,
  "tools": [{"type": "function", "function": {"name": "get_number",
             "description": "Returns the secret number",
             "parameters": {"type": "object", "properties": {}}}}],
  "messages": [
    {"role": "user",
     "content": "Call get_number. Before calling, privately note a random code word in your reasoning. After the call, I will ask you for it."},
    {"role": "assistant", "content": "",
     "reasoning_content": "The user wants me to note a code word privately first. My code word is PELICAN. Now I call get_number.",
     "tool_calls": [{"id": "abc123xyz", "type": "function",
                     "function": {"name": "get_number", "arguments": "{}"}}]},
    {"role": "tool", "tool_call_id": "abc123xyz", "name": "get_number", "content": "391"},
    {"role": "user",
     "content": "What was the code word you noted in your reasoning before the call? Reply with just the word."}
  ]}' | jq '.choices[0].message | {content, reasoning_content}'

The word PELICAN appears nowhere except inside reasoning_content on a replayed assistant turn. If the model answers PELICAN, the reasoning reached it. If the model tells you it never noted a code word, something between you and the model threw it away. Against the unpatched gateway, LiteLLM’s own transform shows what Mistral receives for that assistant turn:

{"role": "assistant", "content": "", "tool_calls": [{"id": "abc123xyz", "type": "function", "function": {"name": "get_number", "arguments": "{}"}}]}

No pelican. The control version of this test, with no reasoning replayed at all, had the model explain in its new reasoning that it had not actually committed to anything in its previous turn, which is exactly right from where it was sitting.

After the patch below:

content=PELICAN
reasoning=We need answer just code word. We previously privately noted PELICAN in reasoning. Need final just PELICAN. Ensure no extra.

I also ran a version with reasoning on a plain text turn rather than a tool call (“I pick 7” in the reasoning, “ready.” as the answer, then “reveal the number”), which exercises the thinking-plus-text shape. Patched: 7, with the model citing “my internal thought process”. Unpatched: a careful explanation that it had never picked a number.

The patch⌗

This is the third function in the runtime patch module from the first post. Same mechanism: the proxy imports the module through litellm_settings.callbacks, module-level code wraps LiteLLM’s own functions in-process, each patch probes itself and raises if it didn’t take, and nothing about the image or the wiring changes. If you deployed the first version, replace the file and roll; the ConfigMap hash does the rest.

There are two hooks this time, because there are two places the thinking dies. The first is the strip function we were already wrapping: before upstream deletes reasoning_content, we move it into Mistral’s chunk format at the head of the content list. The second is the content-list flattener, which the Mistral transform imports by name and calls on every message; it keeps only the text chunks, so we replace it in the Mistral module’s namespace with a version that leaves assistant messages carrying a thinking chunk alone. That second hook is also what makes Vibe-style clients, which already send the right shape, work through the proxy.

import litellm.llms.mistral.chat.transformation as _mistral_module
from litellm.llms.mistral.chat.transformation import MistralConfig

# (the two patches from the first post go here, unchanged except that
#  MISTRAL_EXTRA_FIELDS also lists "reasoning" and "reasoning_details",
#  the OpenRouter-style spellings some clients mirror reasoning_content into)


def _reasoning_text(message) -> str:
    """The reasoning a client replayed on an assistant message, in any of the spellings LiteLLM emits."""
    for key in ("reasoning_content", "reasoning"):
        value = message.get(key)
        if isinstance(value, str) and value.strip():
            return value
    parts = []
    for block in message.get("thinking_blocks") or []:
        if isinstance(block, dict) and block.get("type") == "thinking" and isinstance(block.get("thinking"), str):
            parts.append(block["thinking"])
    return "".join(parts)


def _has_thinking_chunk(content) -> bool:
    return isinstance(content, list) and any(isinstance(c, dict) and c.get("type") == "thinking" for c in content)


def _with_thinking_chunk(message):
    """Copy of an assistant message whose content starts with the Mistral ThinkChunk for its reasoning."""
    text = _reasoning_text(message)
    content = message.get("content")
    if not text or _has_thinking_chunk(content):
        return message
    chunk = {"type": "thinking", "thinking": [{"type": "text", "text": text}]}
    if isinstance(content, list):
        rest = list(content)
    elif isinstance(content, str) and content:
        rest = [{"type": "text", "text": content}]
    else:
        rest = []
    return {**message, "content": [chunk, *rest]}


def _patch_mistral_replay_thinking() -> None:
    upstream_strip = MistralConfig._strip_output_only_fields
    # Both hooks are replaced together; the marker on the strip is the guard for both.
    if getattr(upstream_strip, "_snowcoder_patch", None) == "mistral-replay-thinking":
        return
    # AttributeError here means upstream stopped importing the flattener by this
    # name into its module: re-read _transform_messages in transformation.py.
    upstream_flatten = _mistral_module.handle_messages_with_content_list_to_str_conversion

    def _strip(cls, message):
        if message.get("role") == "assistant":
            message = _with_thinking_chunk(message)  # before upstream deletes reasoning_content
        return upstream_strip(message)

    def _flatten(messages):
        for m in messages:
            if m.get("role") == "assistant" and _has_thinking_chunk(m.get("content")):
                continue  # upstream would keep only the text chunks
            upstream_flatten([m])
        return messages

    _strip._snowcoder_patch = "mistral-replay-thinking"
    _flatten._snowcoder_patch = "mistral-replay-thinking"
    MistralConfig._strip_output_only_fields = classmethod(_strip)
    _mistral_module.handle_messages_with_content_list_to_str_conversion = _flatten

    think = lambda text: {"type": "thinking", "thinking": [{"type": "text", "text": text}]}  # noqa: E731
    probe = [
        {"role": "user", "content": [{"type": "text", "text": "q"}]},
        {
            "role": "assistant",
            "content": "",
            "reasoning_content": "t1",
            "tool_calls": [{"id": "abc123xyz", "type": "function", "function": {"name": "f", "arguments": "{}"}}],
        },
        {"role": "tool", "tool_call_id": "abc123xyz", "name": "f", "content": "r"},
        {"role": "assistant", "content": "a", "reasoning_content": "t2", "provider_specific_fields": {}},
        {"role": "assistant", "content": [think("vibe"), {"type": "text", "text": "v"}]},
        {"role": "assistant", "content": "plain", "thinking_blocks": [{"type": "thinking", "thinking": "t3", "signature": "mistral"}]},
        {"role": "assistant", "content": "none"},
    ]
    got = MistralConfig()._transform_messages([dict(m) for m in probe], "mistral-large-4")
    want_content = ["q", [think("t1")], "r", [think("t2"), {"type": "text", "text": "a"}], [think("vibe"), {"type": "text", "text": "v"}], [think("t3"), {"type": "text", "text": "plain"}], "none"]
    leaked = [k for m in got for k in ("reasoning_content", "reasoning", "thinking_blocks", "provider_specific_fields") if k in m]
    if [m.get("content") for m in got] != want_content or leaked or "tool_calls" not in got[1]:
        raise RuntimeError(f"litellm_patches: Mistral thinking replay patch did not take effect: {got!r}")
    print("litellm_patches: Mistral assistant turns now replay reasoning_content as a thinking chunk", flush=True)


_patch_mistral_replay_thinking()

The probe runs the whole message transform, not just the wrapped function, on a user content list (which must still flatten), a reasoning-plus-tool-call turn, a reasoning-plus-text turn, a Vibe-shaped turn that already carries a chunk, a thinking_blocks turn, and a plain one. If LiteLLM moves either hook in a future release the pod fails to start and the rollout stalls on the old pod, which is the failure mode you want.

Two properties worth knowing. It applies to every mistral/* model, so a conversation that began on another model and switched to Mistral will have that other model’s reasoning replayed too; Mistral tolerates this. And it is not free: every replayed assistant turn now carries its reasoning in the prompt, which is exactly the token-efficiency trade-off Mistral’s documentation describes. On a long agent session that adds up.

Did it matter⌗

I had, as it happens, a Qwen Code job running against Le Chonk while all this was being investigated, and it was not going well: a software development task that had settled into a loop of re-reading files it had already read, working out afresh each turn how its own tools behaved, and making no visible progress. Not an error, just a model that seemed a great deal less capable than its benchmarks suggested. I stopped it, deployed the patch, and started the same task again.

I watched the second run through the gateway’s spend logs rather than the terminal. Seventy-odd calls in the first half hour, none failed, the context growing steadily from twenty-three thousand tokens to a hundred and sixty thousand as it read through the codebase and started writing. The longest calls were genuine generation – sixteen thousand output tokens in one three-minute stretch – rather than stalls. And from the terminal side, the behaviour was different in kind: it opened the right skill files, read the right source, and its visible reasoning followed from the previous step instead of restarting from zero. The file-reading loop did not come back.

I should say plainly that this is one run against one run, and a stuck agent can be stuck for many reasons. But the symptom that went away is precisely the symptom the mechanism predicts. If the model’s conclusion that “I now know how read_file works” lives only in its reasoning, and the reasoning is discarded after every turn, then every turn starts with a model that has forgotten it. Mistral’s documentation says stripping the thinking significantly degrades output quality. I now think they meant it literally.

Is it doing this to your other models?⌗

This was my actual worry at the start, so: the test above works against anything. Replace the model name, run it, see whether you get a pelican.

For my local models the answer was reassuring, and the reason is the provider prefix in the LiteLLM model entry. My llama.cpp and vLLM backends are registered as openai/<name>, which is the transform that leaves reasoning_content alone, and both of them answered PELICAN: llama.cpp with its reasoning-preserve option on a Qwen3 model, and vLLM with its own reasoning parser. So the chain works end to end, LiteLLM is passing it, and the backend templates are rendering it back into the prompt.

The trap is that LiteLLM also offers a hosted_vllm/ prefix, which looks like the obviously correct choice for a vLLM backend, and that transform deletes reasoning_content and thinking_blocks on every assistant turn, exactly as the Mistral one did. If you have a vLLM server behind LiteLLM and you registered it as hosted_vllm/, run the pelican test before assuming your interleaved thinking is interleaving.

Prefix Replayed reasoning Notes
openai/ passed through what my llama.cpp and vLLM entries use
hosted_vllm/ stripped same pattern as Mistral; check with the test
mistral/ stripped, now translated by the patch Mistral’s docs want the chunk
anthropic/, bedrock/, gemini/, ollama/ translated natively the careful path
deepseek/, moonshot/ required; placeholder injected if absent warning logged

Upstream⌗

There was no LiteLLM issue for this, so I have filed one: BerriAI/litellm#45546, with the pelican reproduction and a sketch of the fix, which is a few lines in the Mistral transform. The fix the patch implements is the same one Mistral’s own Vibe client and the Vercel SDK already made, so there is precedent to point at. Until it lands, the module above is the workaround, and it removes itself from consideration the day upstream starts sending the chunk.

The first post ended with the observation that it was about the plumbing rather than the model. That is still true, but this round changed my view of the model somewhat. Judged through a gateway that discards its thinking, Le Chonk looks ordinary. Judged with its thinking intact, it looks like the model the benchmarks describe. It would be worth checking, before concluding anything about any reasoning model you use through a proxy, that the proxy is letting it remember.


Disclosure, in the spirit of transparency about AI use: this post was drafted by Claude, which also did the research it describes – reading the LiteLLM, Vibe and Qwen Code sources, devising the pelican test, writing the patch and watching the spend logs – with me running the clients and the deploys and supplying the opinions; I edited it. The cover is from the same ComfyUI deployment as before; Claude’s prompt asked for a string tied round the cat’s paw as a reminder, and Z-Image Turbo obliged.