~3m13:18Grok 4.7: I Gave It a Broken Championship, a Broken LED, and Coffee
Sep 21, 2026
Read: ~3m · You save: 10 min
Grok 4.7: Testing a Model with a Broken Championship, LED, and Coffee
Testing Grok 4.7: Can AI fix monster truck championship bugs, diagnose a breakdown, and translate 'coffee' into 80 languages? Let's see real tests!
This review tests SpaceX AI's Grok 4.7 model for its effectiveness in solving real-world programming, data analysis, and legal issues. The model is positioned as the most powerful for coding and knowledge work to date, surpassing previous versions in speed and cost.
Technical Specifications and Claims
SpaceX AI claims Grok 4.7 is twice as fast and half the cost of comparable Frontier models. The model was trained using longer reinforcement learning on complex multi-step tasks. Improved self-correction capabilities, long-context processing, and an enhanced safety stack are noted.
Testing on a Real Application
The first stage of testing was to verify Grok 4.7's ability to find and fix errors in a real application. A locally deployed, fully functional application called Dirt Dynasty—a scoring platform for freestyle motocross competitions—was used as the test platform. The application, developed using Docker, PostgreSQL, FastAPI, Nginx, Redis, and other components, had a critical error: the system was only counting scores from three out of four judges, ignoring one score from each participant. This led to incorrect winner determination, placing Gravedigger in first place instead of the actual leader, Max D.
The Grok 4.7 model was tasked with finding and fixing all errors without being directed to a specific problem. After completing the task and restarting the application, it was confirmed that the error was fixed, and Max D secured his rightful first place.
Comparative Benchmarks
SpaceX AI has provided benchmark data demonstrating the performance of Grok 4.7 compared to other models.
- Coding (HumanEval 5.1): Grok 4.7 shows strong results, outperforming some models.
- Terminal Operations and Multi-step Office Tasks: Other models, such as Fable 5.1, lead in these areas.
- Legal and Electrical Engineering Design: Grok 4.7 demonstrates superiority, surpassing all other models in these domains.
- Clinical Diagnosis: The model lags behind GPT 4.6 and Fable 5.1 but generally shows competitive results, positioning itself at the lower end of the price range.
Electrical Engineering Testing
The second test aimed to assess Grok 4.7's electrical engineering capabilities, where the model performed highly on benchmarks. A circuit was presented with a 24V power supply, two parallel resistors, connected in series with a 13-ohm resistor and an LED. A multimeter showed 0V at one point in the circuit, indicating an open circuit or lack of conductivity.
Grok 4.7 successfully diagnosed the fault, identifying that the LED was installed incorrectly (reverse polarity), preventing it from lighting up. The model demonstrated self-correction capabilities by analyzing the image and performing calculations.
Testing in the Legal Domain
To test the model's legal competencies, a generated image of a living room scene with a man and three women, presented as his wives, was used. The task was to analyze the situation from the perspective of legal issues such as marriage validity, cultural specificities, and financial claims, without considering race, age, or appearance.
Grok 4.7 successfully handled the task by analyzing applicable legislation from different jurisdictions, identifying missing facts, and considering issues related to the legality of polygamous marriages in various cultures and religious texts. The model correctly refused to analyze the image as a "soap opera," focusing instead on the legal aspects. Important points were noted, such as kinship ties that could invalidate a marriage, the rights of the presumed spouse, and conflicts of law. The model maintained neutrality, avoiding racial or other biased judgments.
Testing in Programming and Multilingualism
The final test combined two aspects: creating an animation of latte preparation using pure HTML and JavaScript, and checking knowledge of coffee names in 80 languages.
Grok 4.7 successfully generated code for a visually convincing coffee preparation animation, including elements of the machine, milk, and a heart. The animation was deemed comparable in quality to the results of other models.
In the multilingual aspect, the model demonstrated high accuracy, correctly identifying the root of the word "coffee" (e.g., "kahava") in most of the 80 languages and using appropriate scripts. The model also noted languages that borrow words from neighbors and correctly generated a plausible fictional word "corvix" for a hypothetical language.
Conclusion
Grok 4.7 demonstrated impressive results during testing, showcasing its ability to solve complex problems across various domains, including programming, legal analysis, and technical diagnostics. The model confirmed the stated improvements in speed and self-correction, and also showed strong multilingual capabilities.
Introduction to Grok 4.7
Introducing Grok 4.7 as SpaceX AI's most powerful model for coding and knowledge work. Mention of its speed, price, and improved capabilities compared to previous models. Announcement of testing on real-world applications.
- Grok 4.7 is positioned as SpaceX AI's most powerful model for coding and knowledge work.
- The model is claimed to be twice as fast and half the price of comparable Frontier models.
- Long context handling and safety have been improved.
- The model will be tested on real-world application bugs.
Bug Fixing Test in a Monster Truck Championship
Testing Grok 4.7 with the Dirt Dynasty application, a monster truck championship scoring platform. The model must find and fix a bug causing the wrong winner to be determined.
- The Dirt Dynasty application has a bug: one of the four judges' scores is not being accounted for.
- Due to the bug, Gravedigger is leading instead of the actual champion, Max D.
- The application is a complex Docker environment with PostgreSQL, FastAPI, Nginx, and Redis.
- Grok 4.7 is tasked with finding and fixing all bugs without being told the specific problem.
Benchmarks and Performance Comparison
Analysis of Grok 4.7 benchmark results. Comparison with other models (Grok 4.6, Fable 5.1, GPT 5.6 Six) on various tasks, including coding, terminal operations, legal, and medical reasoning. Evaluation of price/quality ratio.
- Grok 4.7 shows impressive benchmark results, especially in legal and engineering tasks.
- The model is the cheapest among those tested.
- In coding and terminal operations, Fable 5.1 outperforms Grok 4.7.
- Grok 4.7 successfully handles legal reasoning, surpassing other models.
Electrical Schematic Diagnosis Test
Testing Grok 4.7's ability to diagnose a fault based on an image of a drawn electrical schematic. The model should identify the cause of zero voltage in the circuit.
- The model's ability to analyze images and diagnose technical issues is being tested.
- The schematic depicts a circuit with a 24V power supply, resistors, an LED, and a multimeter showing 0V.
- Grok 4.7 is expected to identify the cause of the fault (open circuit or incorrect connection).
- The model successfully diagnosed the problem, identifying that the LED was incorrectly connected (reverse polarity).
Test of Legal Analysis for a Complex Scene
Testing Grok 4.7 on a legal task involving the analysis of a polygamy scene. The model should identify legal issues, ignoring race, appearance, and age, focusing on actual legal aspects.
- The scene depicts a man with three wives of different ethnic backgrounds.
- The model's task is to analyze the scene from the perspective of legal issues (marriage validity, conflicts, financial claims).
- The model must ignore racial and age aspects, focusing solely on legal matters.
- Grok 4.7 successfully handled the task, highlighting missing facts, validity questions, laws of different jurisdictions, and cultural nuances.
Animation Coding and Multilingualism Test
Final test: Grok 4.7 should create an HTML/JavaScript animation of latte preparation and correctly name the coffee in 80 languages. Both the model's creativity and linguistic abilities are being tested.
- The task has two parts: creating a coffee preparation animation and translating the word 'coffee' into 80 languages.
- The animation must be visually convincing and created in pure HTML and JavaScript.
- The model must provide the correct names for coffee in their native languages and in the correct transcription.
- Grok 4.7 successfully created the animation and provided correct coffee names, including a plausible made-up word for one language.