~3m9:22
Fahd Mirza

Le chonk: The Mistral Large 4 Thoroughly Tested

Oct 6, 2026

Read: ~3m · You save: 6 min

Le Chonk: Thorough Testing of Mistral Large 4

Mistral Large 4 ('Le Chonk') review: tests on code generation, AI reasoning, security, and multilingualism. Discover the new model's strengths and weaknesses!

Mistral has unveiled its new model, Mistral Large 4, unofficially referred to as ML4 and nicknamed "Le Chunk." This moniker reflects its impressive scale: approximately one trillion parameters, with around 50 billion active at any given time. The model was trained in Mistral's own European data centers. A public preview is currently available via API on Mistral Studio, with open weights expected by the end of this month.

Testing HTML Generation from Image

The first test involved a request to create a single HTML file based on a provided image. A photo of a vertical rotisserie grill with meat and fire was used as input. The model successfully generated an HTML file where the meat layers appear convincing, and a drip pan, metal skewer, and motor are present at the top. However, the image of the fire was identified as a weak point: the flames were small and hidden behind the meat, unlike the bright fire in the original photo. Despite this, the result was rated as "not bad."

Scientific Reasoning Test

The next stage of testing focused on evaluating the model's scientific reasoning capabilities. A task was presented based on the rules depicted in an image: four individuals are touching hot and neutral wires while either standing on insulating mats or on the ground. The model was required to determine which characters would receive an electric shock.

During this test, the model's benchmarking results were also presented.

  • Deep Seek: Tests the model's ability to solve real-world software engineering tasks. ML4 outperforms most presented open models, trailing only Kimiko 3.
  • Terminal Bench: Evaluates the model's performance with command-line instructions. ML4 shows good results compared to most competitors, but GLM 5.3 demonstrates higher performance. The benchmark's fairness is noted, along with the superiority of Chinese models over Mistral and ML4 in this test.
  • Cyber Index: Measures the model's effectiveness in finding and fixing security vulnerabilities in real software. ML4 holds a leading position, sharing it with GLM 5.3 Flash.
  • Business Workflows: Tests the model's ability to execute business processes, including working with email, spreadsheets, and messengers. ML4 lags behind the most powerful open models, trailing GLM 5.3.
  • Finance Agent: Evaluates the model's performance in researching and analyzing company financial statements. ML4 is positioned at the top of the list, slightly behind GLM and marginally ahead of GPT-6 Astra.

After completing the scientific reasoning test, the model provided an answer stating that characters A and D would receive an electric shock. A is touching a hot wire connected to the ground, and D is touching both hot and neutral wires simultaneously. The model correctly identified that characters B and C are in safe conditions (B is insulated, C has the same potential as the ground). ML4 demonstrated the ability to read rules from an image, perform step-by-step reasoning, and handle complex cases.

Vulnerability Discovery in Code Testing

Another claimed advantage of the model is its ability to find vulnerabilities in code. As an example, a Flask client portal with an SQLite database containing user and account information was presented. When attempting to access an endpoint, the user Alice requested account 102, which belonged to Bob. The system returned complete information about Bob's account, which is a broken access control error, as Alice should not have access to other users' data.

The ML4 model, working in conjunction with the Hermes agent, was tasked with finding this vulnerability and suggesting fixes. The model quickly identified an Insecure Direct Object Reference (IDOR) vulnerability, noting that the delete endpoint checks ownership, while the retrieval endpoint does not. Additionally, ML4 discovered other vulnerabilities: plaintext passwords, a hardcoded session key, and a cookie with an active session token. The model provided correct code for the fixes and suggested applying these fixes.

"Reading Between the Lines" Test

To check the model's ability to interpret implicit information, a screenshot of a WhatsApp conversation was used. In the message, an employee is communicating with their boss, with the message timestamps not matching the context. The model's task was to understand what was actually happening. This test checks the ability to reason in social situations, not just to read text.

ML4 successfully completed this task, correctly interpreting the messages as a "trigger," understanding the work slang ("left me hanging"), and inferring that the boss's wife misinterpreted the message as romantic. The model constructed a clear chain of reasoning from the ambiguous text to the boss's panicked response, which fully met the test's objective. The result was rated as "10 out of 10."

Multilingualism Test

The final test involved the front page of a Tamil newspaper. The model was tasked with reading the headline, translating it into 75 languages and scripts, including Mandarin, Arabic, Esperanto, and fictional languages. This test assesses the model's ability to handle texts in various scripts, not just the mainstream languages for which models are typically optimized.

The model successfully translated the headline into approximately 25 languages, including correct script usage for Tamil, Arabic, Japanese, and Thai. However, when dealing with low-resource languages like Gujarati, the model began to make errors, and further translations were inaccurate. Thus, multilingualism, especially for low-resource languages, is not ML4's strong suit.

Introducing Mistral Large 4 ("Le Chonk")

Introducing the Mistral Large 4 model, unofficially nicknamed "Le Chonk." Mention of its size (trillion parameters) and availability (public preview API, open weights expected).

  • The Mistral Large 4 model is unofficially nicknamed "Le Chonk."
  • Total number of parameters is around one trillion.
  • Approximately 50 billion active parameters.
  • Trained in Mistral's European data centers.
  • Available in public preview via the Mistral Studio API.
  • Open weights are expected by the end of the month.

Test 1: HTML Generation from Image

Test of generating an HTML file from a doner image and text description. The model successfully created HTML, but with some shortcomings in the fire image.

  • Test: creating an HTML file from an image and prompt.
  • Used an image of a vertical grill with meat.
  • The model generated an HTML file.
  • The image of the meat and grill details turned out convincing.
  • Weaknesses noted: insufficiently realistic fire.

Test 2: Scientific Reasoning

Scientific reasoning test: determining who will get shocked when touching wires. The model correctly identified the victims (A and D) and explained the reasons.

  • Test: Scientific reasoning based on rules from the image.
  • Task: Determine who will get shocked by hot/neutral wires.
  • Model correctly identified that A and D will get shocked.
  • Explanation: A is grounded, D touches both wires.
  • Model considered B's insulation and C's equal potential with the floor.

Performance Benchmarking

Analysis of Mistral Large 4 benchmarking results across various tasks: Deep Seek (Software Engineering), Terminal Bench (Terminal), Cyber Index (Security), Business Workflows (Business Processes), Finance Agent (Finance).

  • Deep Seek: ML4 outperforms most open models, trailing Kimiko 3.
  • Terminal Bench: ML4 performs well, but GLM 5.3 is ahead.
  • Cyber Index: ML4 is at the top, on par with GLM 5.3 Flash.
  • Business Workflows: ML4 lags behind the strongest open models, trailing GLM 5.3.
  • Finance Agent: ML4 slightly trails GLM and leads GPT-6 Astra.

Test 3: Code Vulnerability Scan

Test for finding vulnerabilities in code. The model detected a Broken Access Control (IDOR) vulnerability, as well as issues with plaintext passwords and session tokens.

  • Test: finding vulnerabilities in Flask application code.
  • IDOR (Insecure Direct Object Reference) vulnerability detected.
  • Vulnerability: the get endpoint does not check data ownership.
  • Additionally found: plaintext passwords, hardcoded session key, cookie file with token.
  • Model suggested code fixes.

Test 4: Reading Between the Lines (Social Reasoning)

Test of understanding social context from a WhatsApp screenshot. The model correctly interpreted implicit messages, sarcasm, and the boss's reaction.

  • Test: analysis of social interaction from a WhatsApp screenshot.
  • Task: understand the hidden meaning of messages between an employee and a boss.
  • Model correctly understood sarcasm ('left me hanging').
  • Model correctly interpreted the boss's reaction as panic.
  • Rating: 10 out of 10 for social reasoning.

Test 5: Multilingualism

Multilingual test: translating a newspaper headline into 75 languages. The model successfully handled major languages and scripts but showed weakness in low-resource languages.

  • Test: translating a Tamil newspaper headline into multiple languages.
  • Goal: testing the processing of texts in different scripts and languages.
  • Approximately 25 languages were successfully translated, including major ones (Arabic, Japanese, Thai).
  • Problems arose with low-resource languages (e.g., Gujarati).
  • Conclusion: multilingualism is not the model's strongest suit for low-resource languages.